Speech runs on the server: audio and transcripts are not sent anywhere else. Find speech model ids in GET /v1/models/catalog.
Text-to-speech
POST /v1/audio/speech returns audio, streamed as it is generated.
with client.audio.speech.with_streaming_response.create(
model="<tts-model-id>", voice="<voice>", input="Your meeting starts in five minutes.",
response_format="mp3") as response:
response.stream_to_file("reminder.mp3")
| Parameter | Notes |
|---|---|
model, input | Required. |
voice | A voice the model offers. |
response_format | mp3 (default), opus, aac, flac, wav, or pcm (raw 16-bit, 24 kHz, mono). |
speed | Speaking rate; 1.0 is normal. |
Transcription
POST /v1/audio/transcriptions takes a multipart/form-data upload.
curl http://<server>/v1/audio/transcriptions -H "Authorization: Bearer $KEY" \
-F [email protected] -F model=<stt-model-id> -F response_format=srt -o meeting.srt
| Field | Notes |
|---|---|
file, model | Required. Local models take WAV (PCM or float, any sample rate); see below for MP3 and M4A. |
language | Two-letter code; omit to detect. |
prompt | Words to expect (names, jargon). |
temperature | Sampling temperature. |
response_format | json (default), text, verbose_json (with segments and timings), srt, vtt. |
stream | true for Server-Sent Events: transcript.text.delta chunks, then transcript.text.done. |
Translation
POST /v1/audio/translations takes the same fields and returns English text from speech in any language the model knows.
Live transcription
GET /v1/audio/transcriptions/stream is a WebSocket for captions while people speak.
- Open
ws://<server>/v1/audio/transcriptions/stream?model=<runtime/family/variant>&language=en(orwss://with TLS). The model must be given in theruntime/family/variantform. - Send audio as binary frames: 16 kHz, mono, 16-bit PCM.
- Receive text frames with
transcript.text.deltaand, at the end,transcript.text.done. - Close the socket to say the audio has ended; the server sends the final text and closes.
Authenticate with the Authorization header where your client allows it. Browsers cannot set headers on a WebSocket, so ?key=<api-key> is accepted on this endpoint only, and never written to logs. A page from another origin can open the socket only if that origin is in AISUITE_CORS_ORIGINS. A session lasts at most two hours.
Voice cloning
POST /v1/audio/clone speaks text in the voice of a short reference recording.
| Field | Notes |
|---|---|
file | Required. A clean recording of the voice, a few seconds to a minute. |
input | Required. The text to speak. |
engine | chatterbox (default), openvoice, zonos or metavoice, where installed. |
language | Default en. |
speed, stream | Rate, and streamed output. |
The answer is WAV audio. The reference recording is used for this request only and is never stored or logged. Clone only voices you have permission to use.
Questions
Which audio formats can I upload? +
Local speech models take WAV files (16-bit PCM or float, mono or stereo, any sample rate; the server converts them). Convert compressed recordings first, for example ffmpeg -i meeting.m4a -ar 16000 -ac 1 meeting.wav. A cloud speech provider, if an operator has configured one, accepts the formats that provider supports.
Do I need a GPU for speech? +
Transcription and text-to-speech run on a CPU, more slowly. A GPU makes live captions and long recordings practical.