Text to speech over a WebSocket
Open one WebSocket, send text as it is produced, and receive each response's audio as it is synthesized.
The speech WebSocket lets you synthesize many utterances over one connection. You send text as conversation items, ask for a response, and the audio streams back as base64 PCM16 deltas. A voice agent uses it to speak each sentence as soon as the language model produces it.
The events follow the OpenAI Realtime shapes for text input and audio output.
For a single clip from one request, use POST /v1/audio/speech.
Speak a sentence
import asyncio, base64, json, os
import websockets
async def main():
url = "wss://api.sprag.ai/v1/audio/speech?model=chorus"
headers = {"Authorization": f"Bearer {os.environ['SPRAG_API_KEY']}"}
async with websockets.connect(url, additional_headers=headers) as ws:
await ws.send(json.dumps({
"type": "session.update",
"session": {"audio": {"output": {"voice": "serena"}}},
}))
await ws.send(json.dumps({
"type": "conversation.item.create",
"item": {
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "The quarterly numbers are in."}],
},
}))
await ws.send(json.dumps({"type": "response.create"}))
pcm = bytearray()
async for message in ws:
event = json.loads(message)
if event["type"] == "response.output_audio.delta":
pcm += base64.b64decode(event["delta"])
elif event["type"] == "response.done":
break
print(f"{len(pcm)} bytes of 24 kHz PCM16")
asyncio.run(main())The first event is session.created, which states the output format and the
voice the session opens in:
{
"type": "session.created",
"session": {
"object": "realtime.session",
"type": "realtime",
"model": "chorus",
"output_modalities": ["audio"],
"audio": {"output": {"format": {"type": "audio/pcm", "rate": 24000}, "voice": "dominic"}},
"instructions": "",
"expires_at": 1791227832
}
}How it works
Text you send is held until you call response.create. Each
response.create turns the text sent since the previous one into a single
utterance and streams its audio. Responses are spoken one at a time; a
response.create sent while one is speaking waits for it to finish.
The voice, speed, and instructions are fixed once the first response is created. To speak in another voice, open a new session.
Connect
wss://api.sprag.ai/v1/audio/speech?model=chorusmodel takes a text-to-speech model id from Models.
Authenticate the upgrade with your API key in the bearer header:
Authorization: Bearer $SPRAG_API_KEYA client that cannot set headers on a WebSocket upgrade, such as the browser
WebSocket API, passes the key as a subprotocol pair instead: sprag-api-key
followed by the key.
new WebSocket("wss://api.sprag.ai/v1/audio/speech?model=chorus", [
"sprag-api-key",
spragApiKey,
]);Configure the session
Send session.update before the first response.create. The server answers
each accepted update with session.updated.
{
"type": "session.update",
"session": {
"audio": {"output": {"voice": "serena"}},
"instructions": "Warm and friendly, moderate pacing."
}
}| Field | Required | Type | Notes |
|---|---|---|---|
session.audio.output.voice | No | string | A preset id, or the id of a voice your organization created. Defaults to the voice in session.created. |
session.instructions | No | string | Style guidance. Accepted for voices whose model supports it. |
Voices lists the presets and how to create your own.
Send text
Send text as a conversation.item.create carrying an input_text part, then
response.create to speak it:
{
"type": "conversation.item.create",
"item": {
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "The quarterly numbers are in."}]
}
}{"type": "response.create"}The server confirms each item with conversation.item.added and
conversation.item.done. To speak a sentence as soon as it is ready, send one
item and one response.create per sentence.
Read the audio
| Event | Meaning |
|---|---|
response.created | A response started. |
response.output_audio_transcript.delta | The text the response speaks. |
response.output_audio.delta | A chunk of audio: base64 PCM16, mono, 24 kHz. |
response.output_audio.done | The response's audio is complete. |
response.done | The response finished. response.status is completed, cancelled, or failed, and response.usage reports the audio produced. |
Cancel a response
{"type": "response.cancel"}The active response ends with status cancelled, and responses waiting
behind it are cancelled with it. A cancel with no response active returns the
error response_cancel_not_active.
Constraints
- An utterance holds at most 8,192 characters.
- A session accepts at most 8,192 characters of text per minute.
- A session runs for at most one hour.
session.createdcarries the end asexpires_at. - The session takes text only. Audio input events and binary frames return an
error.
Errors
The upgrade fails with an HTTP status and a JSON error body before the socket opens:
| Status | Meaning |
|---|---|
400 | model is missing, or the model does not serve speech. |
401 | The API key is missing or invalid. |
404 | The model does not exist. |
429 | The rate limit is spent. |
During a session, a refused event returns an error event and the session
continues. The code names the reason:
| Code | Meaning |
|---|---|
unknown_voice | The model has no voice with that id. |
unsupported_parameter | The voice's model does not accept the field, such as instructions on a cloned voice. |
session_config_locked | The voice or instructions changed after the first response was created. |
input_too_long | The utterance would exceed 8,192 characters. |
rate_limit_exceeded | The session sent more than 8,192 characters in the last minute. |
response_queue_full | Too many responses are waiting behind the one speaking. |
model_not_running, model_at_capacity | No replica could take the session when its first response was created. The socket then closes with code 1013; retry on a new session. |
Next steps
To create a voice of your own, see Voice cloning and Voice design. To transcribe live audio over a WebSocket, see Speech to text in realtime.