Speech runs on the server: audio and transcripts are not sent anywhere else. Find speech model ids in GET /v1/models/catalog.

Text-to-speech

POST /v1/audio/speech returns audio, streamed as it is generated.

with client.audio.speech.with_streaming_response.create(
        model="<tts-model-id>", voice="<voice>", input="Your meeting starts in five minutes.",
        response_format="mp3") as response:
    response.stream_to_file("reminder.mp3")
ParameterNotes
model, inputRequired.
voiceA voice the model offers.
response_formatmp3 (default), opus, aac, flac, wav, or pcm (raw 16-bit, 24 kHz, mono).
speedSpeaking rate; 1.0 is normal.

Transcription

POST /v1/audio/transcriptions takes a multipart/form-data upload.

curl http://<server>/v1/audio/transcriptions -H "Authorization: Bearer $KEY" \
  -F [email protected] -F model=<stt-model-id> -F response_format=srt -o meeting.srt
FieldNotes
file, modelRequired. Local models take WAV (PCM or float, any sample rate); see below for MP3 and M4A.
languageTwo-letter code; omit to detect.
promptWords to expect (names, jargon).
temperatureSampling temperature.
response_formatjson (default), text, verbose_json (with segments and timings), srt, vtt.
streamtrue for Server-Sent Events: transcript.text.delta chunks, then transcript.text.done.

Translation

POST /v1/audio/translations takes the same fields and returns English text from speech in any language the model knows.

Live transcription

GET /v1/audio/transcriptions/stream is a WebSocket for captions while people speak.

  1. Open ws://<server>/v1/audio/transcriptions/stream?model=<runtime/family/variant>&language=en (or wss:// with TLS). The model must be given in the runtime/family/variant form.
  2. Send audio as binary frames: 16 kHz, mono, 16-bit PCM.
  3. Receive text frames with transcript.text.delta and, at the end, transcript.text.done.
  4. Close the socket to say the audio has ended; the server sends the final text and closes.

Authenticate with the Authorization header where your client allows it. Browsers cannot set headers on a WebSocket, so ?key=<api-key> is accepted on this endpoint only, and never written to logs. A page from another origin can open the socket only if that origin is in AISUITE_CORS_ORIGINS. A session lasts at most two hours.

Voice cloning

POST /v1/audio/clone speaks text in the voice of a short reference recording.

FieldNotes
fileRequired. A clean recording of the voice, a few seconds to a minute.
inputRequired. The text to speak.
enginechatterbox (default), openvoice, zonos or metavoice, where installed.
languageDefault en.
speed, streamRate, and streamed output.

The answer is WAV audio. The reference recording is used for this request only and is never stored or logged. Clone only voices you have permission to use.

Questions

Which audio formats can I upload? +

Local speech models take WAV files (16-bit PCM or float, mono or stereo, any sample rate; the server converts them). Convert compressed recordings first, for example ffmpeg -i meeting.m4a -ar 16000 -ac 1 meeting.wav. A cloud speech provider, if an operator has configured one, accepts the formats that provider supports.

Do I need a GPU for speech? +

Transcription and text-to-speech run on a CPU, more slowly. A GPU makes live captions and long recordings practical.