Sonar logo

Sonar

A self-hosted LLM inference server, formerly Aphrodite Engine, that serves Hugging Face models through OpenAI-compatible APIs.

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

About Sonar

Sonar, previously known as Aphrodite Engine, is an inference engine for Hugging Face-compatible language and multimodal models. It is based on vLLM and adds more model and quantization formats, sampling methods, kernels, platforms and deployment features. It is built to serve models to many users at once, with continuous batching, paged KV-cache management, prefix caching, speculative decoding and tensor, pipeline, data and expert parallelism.

The server listens locally and provides OpenAI-compatible APIs, health checks, metrics and an OpenAPI schema. An installer covers Linux on NVIDIA, AMD or CPU, Linux Arm64 and Apple silicon Macs.

Strengths

  • OpenAI-compatible API with streaming, embeddings and tool calls
  • Continuous batching and several forms of parallelism for many concurrent users
  • Wide quantization and model format support beyond vLLM
  • Supports NVIDIA, AMD, Intel, CPU, Apple silicon and TPU

Limitations

  • Aimed at server deployments rather than casual desktop chat
  • Recently renamed from Aphrodite Engine, so older guides use the old name
  • Best performance needs a capable GPU

Details

Pricing
FreeFree and open source under the AGPL-3.0 licence.
License
AGPL-3.0
Developer
dphnAI
Platforms
macOS, Linux, Self-hosted, Command line
How it runs
Downloadable app, Self-hosted
Account
Not required
Best suited for
Serving local or open-weight language models to many users over an API
Last verified
Added

Alternatives to Sonar

Compare all

Software that can replace Sonar for an important use case, and what changes if you switch.

  • vLLM

    High-throughput, memory-efficient inference and serving engine for LLMs.

    vLLM is the Apache-2.0 licensed high-throughput serving engine for Linux, moving from Sonar's AGPL-3.0 license, with narrower quantization and hardware support than Sonar lists.

  • SGLang

    A serving framework for running large language models and multimodal models on your own GPUs.

    SGLang is an open-source serving framework built for high-throughput language and multimodal model serving on your own GPUs, with Docker files and benchmark tooling included.

  • TabbyAPI

    A lightweight, OpenAI-compatible API server for running ExLlama-format language models on your own hardware.

    TabbyAPI is a lightweight OpenAI-compatible server that only serves ExLlama-format models, adds LoRA support and sampler overrides, and has no built-in chat interface.

  • Xinference

    An open-source inference server for running language, speech and multimodal models through one API.

    Xinference serves language, speech and multimodal models from one GPT-compatible API and runs on a laptop or scales to servers, rather than Sonar's many-user batching focus.

  • OpenLLM

    A command-line tool that serves open-source LLMs as OpenAI-compatible APIs on your own hardware.

    OpenLLM is an Apache-2.0 command-line tool that starts an OpenAI-compatible server in one command, with a built-in chat UI and a wide catalogue of open models.

  • LocalAI

    Serve language, image and speech models on your own hardware.

    LocalAI is an MIT-licensed server that also handles image and speech models, offering compatible APIs for several common AI services, though you supply model files and resources.

  • mistral.rs

    A Rust inference engine for running text, vision and multimodal models locally with an OpenAI-compatible server.

    mistral.rs is an MIT-licensed Rust engine with an OpenAI and Anthropic compatible server, multimodal input and built-in GGUF loading, aimed at fast scriptable local use.

  • NVIDIA Triton Inference Server

    An open-source inference server for deploying machine learning models from many frameworks on GPUs and CPUs.

    NVIDIA Triton Inference Server serves models from many ML frameworks on GPUs and CPUs in the cloud or at the edge, requiring more setup and tuning effort.

Sonar as an alternative

Listings that name Sonar as an alternative.

  • OpenVINO Model Server

    A self-hosted inference server for serving AI models that have been optimized with OpenVINO.

    Sonar serves Hugging Face language models through OpenAI-compatible APIs on NVIDIA, AMD, Intel, CPU and Apple hardware, under the AGPL-3.0 license.

Similar software

Related functionality, not necessarily a direct replacement.

Report a wrong fact or a dead link on this listing