<!-- generated: do not edit. source: content/docs/tts/websocket.md -->

# Text to speech over a WebSocket

The speech WebSocket lets you synthesize many utterances over one connection.
You send text as conversation items, ask for a response, and the audio streams
back as base64 PCM16 deltas. A voice agent uses it to speak each sentence as
soon as the language model produces it.

The events follow the OpenAI Realtime shapes for text input and audio output.
For a single clip from one request, use [`POST /v1/audio/speech`](/docs/tts/api).

## Speak a sentence

```python
import asyncio, base64, json, os
import websockets

async def main():
    url = "wss://api.sprag.ai/v1/audio/speech?model=chorus"
    headers = {"Authorization": f"Bearer {os.environ['SPRAG_API_KEY']}"}
    async with websockets.connect(url, additional_headers=headers) as ws:
        await ws.send(json.dumps({
            "type": "session.update",
            "session": {"audio": {"output": {"voice": "serena"}}},
        }))
        await ws.send(json.dumps({
            "type": "conversation.item.create",
            "item": {
                "type": "message",
                "role": "user",
                "content": [{"type": "input_text", "text": "The quarterly numbers are in."}],
            },
        }))
        await ws.send(json.dumps({"type": "response.create"}))

        pcm = bytearray()
        async for message in ws:
            event = json.loads(message)
            if event["type"] == "response.output_audio.delta":
                pcm += base64.b64decode(event["delta"])
            elif event["type"] == "response.done":
                break
        print(f"{len(pcm)} bytes of 24 kHz PCM16")

asyncio.run(main())
```

The first event is `session.created`, which states the output format and the
voice the session opens in:

```json
{
  "type": "session.created",
  "session": {
    "object": "realtime.session",
    "type": "realtime",
    "model": "chorus",
    "output_modalities": ["audio"],
    "audio": {"output": {"format": {"type": "audio/pcm", "rate": 24000}, "voice": "dominic"}},
    "instructions": "",
    "expires_at": 1791227832
  }
}
```

## How it works

Text you send is held until you call `response.create`. Each
`response.create` turns the text sent since the previous one into a single
utterance and streams its audio. Responses are spoken one at a time; a
`response.create` sent while one is speaking waits for it to finish.

The voice, speed, and instructions are fixed once the first response is
created. To speak in another voice, open a new session.

## Connect

```txt
wss://api.sprag.ai/v1/audio/speech?model=chorus
```

`model` takes a text-to-speech model id from [Models](/docs/tts/models).

Authenticate the upgrade with your API key in the bearer header:

```txt
Authorization: Bearer $SPRAG_API_KEY
```

A client that cannot set headers on a WebSocket upgrade, such as the browser
`WebSocket` API, passes the key as a subprotocol pair instead: `sprag-api-key`
followed by the key.

```javascript
new WebSocket("wss://api.sprag.ai/v1/audio/speech?model=chorus", [
  "sprag-api-key",
  spragApiKey,
]);
```

## Configure the session

Send `session.update` before the first `response.create`. The server answers
each accepted update with `session.updated`.

```json
{
  "type": "session.update",
  "session": {
    "audio": {"output": {"voice": "serena"}},
    "instructions": "Warm and friendly, moderate pacing."
  }
}
```

| Field | Required | Type | Notes |
| --- | --- | --- | --- |
| `session.audio.output.voice` | No | string | A preset id, or the id of a voice your organization created. Defaults to the voice in `session.created`. |
| `session.instructions` | No | string | Style guidance. Accepted for voices whose model supports it. |

[Voices](/docs/tts/voices) lists the presets and how to create your own.

## Send text

Send text as a `conversation.item.create` carrying an `input_text` part, then
`response.create` to speak it:

```json
{
  "type": "conversation.item.create",
  "item": {
    "type": "message",
    "role": "user",
    "content": [{"type": "input_text", "text": "The quarterly numbers are in."}]
  }
}
```

```json
{"type": "response.create"}
```

The server confirms each item with `conversation.item.added` and
`conversation.item.done`. To speak a sentence as soon as it is ready, send one
item and one `response.create` per sentence.

## Read the audio

| Event | Meaning |
| --- | --- |
| `response.created` | A response started. |
| `response.output_audio_transcript.delta` | The text the response speaks. |
| `response.output_audio.delta` | A chunk of audio: base64 PCM16, mono, 24 kHz. |
| `response.output_audio.done` | The response's audio is complete. |
| `response.done` | The response finished. `response.status` is `completed`, `cancelled`, or `failed`, and `response.usage` reports the audio produced. |

## Cancel a response

```json
{"type": "response.cancel"}
```

The active response ends with `status` `cancelled`, and responses waiting
behind it are cancelled with it. A cancel with no response active returns the
error `response_cancel_not_active`.

## Constraints

- An utterance holds at most 8,192 characters.
- A session accepts at most 8,192 characters of text per minute.
- A session runs for at most one hour. `session.created` carries the end as `expires_at`.
- The session takes text only. Audio input events and binary frames return an `error`.

## Errors

The upgrade fails with an HTTP status and a JSON error body before the socket
opens:

| Status | Meaning |
| --- | --- |
| `400` | `model` is missing, or the model does not serve speech. |
| `401` | The API key is missing or invalid. |
| `404` | The model does not exist. |
| `429` | The rate limit is spent. |

During a session, a refused event returns an `error` event and the session
continues. The `code` names the reason:

| Code | Meaning |
| --- | --- |
| `unknown_voice` | The model has no voice with that id. |
| `unsupported_parameter` | The voice's model does not accept the field, such as `instructions` on a cloned voice. |
| `session_config_locked` | The voice or instructions changed after the first response was created. |
| `input_too_long` | The utterance would exceed 8,192 characters. |
| `rate_limit_exceeded` | The session sent more than 8,192 characters in the last minute. |
| `response_queue_full` | Too many responses are waiting behind the one speaking. |
| `model_not_running`, `model_at_capacity` | No replica could take the session when its first response was created. The socket then closes with code `1013`; retry on a new session. |

## Next steps

To create a voice of your own, see [Voice cloning](/docs/tts/voices/cloning)
and [Voice design](/docs/tts/voices/design). To transcribe live audio over a
WebSocket, see [Speech to text in realtime](/docs/stt/api/realtime).
