Skip to content

Voice agents that listen, think, and talk back in one API.

Native speech-to-speech agents, plus the full STT → LLM → TTS pipeline when you need control at every step.

OpenAI-compatible APISprag APISelf-hosted

Talk to Symphony

Time to first agent audio

  • Sprag1,033 ms
  • OpenAI1,458 ms
  • Cartesia1,896 ms
  • AssemblyAI3,107 ms
  • Deepgram3,299 ms

Four ways to make a voice.

Pick one and try it here. Everything on this page runs on the same account you sign up for, and the voices carry between all four.

153 / 500

Every voice speaks every language.

62 preset voices, each with its own origin and accent. Any of them can be asked for any of the 10 output languages below, and fidelity is strongest in the one they came from.

Output languages: English, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish.

Eight of 62. Every voice you clone or design joins them.

Browse the full library

Enterprise-grade infrastructure, built for speed.

Sprag serves its own models on open-source infrastructure rather than reselling somebody else's endpoint. That is where the latency goes, and it is why our audio engineers can tune a voice for you when a stock one will not do.

Time to first agent audioMethodology

Time to first agent audio by provider, lowest first: Sprag 1033 milliseconds (lowest), OpenAI 1458 milliseconds, Cartesia 1896 milliseconds, AssemblyAI 3107 milliseconds, Deepgram 3299 milliseconds.

1.18s p50 end-to-end turn latency on agentic_basic, n=30, against OpenAI at 1.76s and Cartesia at 2.24s. Speech in, speech out through a single model.

See the methodology

Choose your architecture

Both run on the same API. Switch per use case, not per vendor.

Speech-to-speech (Symphony)

Need the lowest latency and the most natural turn-taking? Use Symphony, one model end to end.

  • ~1 s to first agent audio
  • Realtime tool calling
  • Natural interruptions and turn-taking
Explore agents

Cascading pipeline (Rhythm → LLM → Chorus)

Need guardrails at the text layer, a specific LLM, or per-step logs? Use the cascade, with LLMs we serve and tune for speed, or your own.

  • Sprag-served LLMs, tuned to cut turn latency
  • Swap any stage, including your own LLM
  • Full transcripts and per-step observability
Explore the pipeline
vLLM-Omni

Founded on Open-Source

The inference engine built for every modality.

−91.4%

reduction in job completion time arXiv↗

Text
Image
Audio
Video

Omni-modal

Every modality. One surface.

Text, image, audio, and video inference in a single API layer.

Any architecture

AR, DiT, and parallel generation.

Autoregressive, diffusion transformers, and non-autoregressive models in one engine.

OpenAI + ComfyUI

Works with the client you have.

Fully OpenAI-compatible API server. Native ComfyUI support.

KV cache efficiency

vLLM memory management, extended.

State-of-the-art KV cache from vLLM core, applied to multimodal workloads.

Make your first request.