Skip to content

Text to speech over a WebSocket

Open one WebSocket, send text as it is produced, and receive each response's audio as it is synthesized.

The speech WebSocket lets you synthesize many utterances over one connection. You send text as conversation items, ask for a response, and the audio streams back as base64 PCM16 deltas. A voice agent uses it to speak each sentence as soon as the language model produces it.

The events follow the OpenAI Realtime shapes for text input and audio output. For a single clip from one request, use POST /v1/audio/speech.

Speak a sentence

import asyncio, base64, json, os
import websockets

async def main():
    url = "wss://api.sprag.ai/v1/audio/speech?model=chorus"
    headers = {"Authorization": f"Bearer {os.environ['SPRAG_API_KEY']}"}
    async with websockets.connect(url, additional_headers=headers) as ws:
        await ws.send(json.dumps({
            "type": "session.update",
            "session": {"audio": {"output": {"voice": "serena"}}},
        }))
        await ws.send(json.dumps({
            "type": "conversation.item.create",
            "item": {
                "type": "message",
                "role": "user",
                "content": [{"type": "input_text", "text": "The quarterly numbers are in."}],
            },
        }))
        await ws.send(json.dumps({"type": "response.create"}))

        pcm = bytearray()
        async for message in ws:
            event = json.loads(message)
            if event["type"] == "response.output_audio.delta":
                pcm += base64.b64decode(event["delta"])
            elif event["type"] == "response.done":
                break
        print(f"{len(pcm)} bytes of 24 kHz PCM16")

asyncio.run(main())

The first event is session.created, which states the output format and the voice the session opens in:

{
  "type": "session.created",
  "session": {
    "object": "realtime.session",
    "type": "realtime",
    "model": "chorus",
    "output_modalities": ["audio"],
    "audio": {"output": {"format": {"type": "audio/pcm", "rate": 24000}, "voice": "dominic"}},
    "instructions": "",
    "expires_at": 1791227832
  }
}

How it works

Text you send is held until you call response.create. Each response.create turns the text sent since the previous one into a single utterance and streams its audio. Responses are spoken one at a time; a response.create sent while one is speaking waits for it to finish.

The voice, speed, and instructions are fixed once the first response is created. To speak in another voice, open a new session.

Connect

wss://api.sprag.ai/v1/audio/speech?model=chorus

model takes a text-to-speech model id from Models.

Authenticate the upgrade with your API key in the bearer header:

Authorization: Bearer $SPRAG_API_KEY

A client that cannot set headers on a WebSocket upgrade, such as the browser WebSocket API, passes the key as a subprotocol pair instead: sprag-api-key followed by the key.

new WebSocket("wss://api.sprag.ai/v1/audio/speech?model=chorus", [
  "sprag-api-key",
  spragApiKey,
]);

Configure the session

Send session.update before the first response.create. The server answers each accepted update with session.updated.

{
  "type": "session.update",
  "session": {
    "audio": {"output": {"voice": "serena"}},
    "instructions": "Warm and friendly, moderate pacing."
  }
}
FieldRequiredTypeNotes
session.audio.output.voiceNostringA preset id, or the id of a voice your organization created. Defaults to the voice in session.created.
session.instructionsNostringStyle guidance. Accepted for voices whose model supports it.

Voices lists the presets and how to create your own.

Send text

Send text as a conversation.item.create carrying an input_text part, then response.create to speak it:

{
  "type": "conversation.item.create",
  "item": {
    "type": "message",
    "role": "user",
    "content": [{"type": "input_text", "text": "The quarterly numbers are in."}]
  }
}
{"type": "response.create"}

The server confirms each item with conversation.item.added and conversation.item.done. To speak a sentence as soon as it is ready, send one item and one response.create per sentence.

Read the audio

EventMeaning
response.createdA response started.
response.output_audio_transcript.deltaThe text the response speaks.
response.output_audio.deltaA chunk of audio: base64 PCM16, mono, 24 kHz.
response.output_audio.doneThe response's audio is complete.
response.doneThe response finished. response.status is completed, cancelled, or failed, and response.usage reports the audio produced.

Cancel a response

{"type": "response.cancel"}

The active response ends with status cancelled, and responses waiting behind it are cancelled with it. A cancel with no response active returns the error response_cancel_not_active.

Constraints

  • An utterance holds at most 8,192 characters.
  • A session accepts at most 8,192 characters of text per minute.
  • A session runs for at most one hour. session.created carries the end as expires_at.
  • The session takes text only. Audio input events and binary frames return an error.

Errors

The upgrade fails with an HTTP status and a JSON error body before the socket opens:

StatusMeaning
400model is missing, or the model does not serve speech.
401The API key is missing or invalid.
404The model does not exist.
429The rate limit is spent.

During a session, a refused event returns an error event and the session continues. The code names the reason:

CodeMeaning
unknown_voiceThe model has no voice with that id.
unsupported_parameterThe voice's model does not accept the field, such as instructions on a cloned voice.
session_config_lockedThe voice or instructions changed after the first response was created.
input_too_longThe utterance would exceed 8,192 characters.
rate_limit_exceededThe session sent more than 8,192 characters in the last minute.
response_queue_fullToo many responses are waiting behind the one speaking.
model_not_running, model_at_capacityNo replica could take the session when its first response was created. The socket then closes with code 1013; retry on a new session.

Next steps

To create a voice of your own, see Voice cloning and Voice design. To transcribe live audio over a WebSocket, see Speech to text in realtime.