Skip to main content

Text to Speech

Convert text to natural-sounding speech — Indic, English and multilingual TTS models behind one OpenAI-compatible endpoint.

Overview

The Text to Speech API converts text into audio with any of 9 TTS models. The default, bulbul:v3, has 37 voices across 11 Indian languages.

Four models are on the free tier: bulbul:v3, aura-2-en, aura-2-es and melotts.

Endpoint: POST /v1/audio/speech

  1. 1

    Your app

    Send text + a and to

  2. 2

    Your chosen model

    Synthesize speech in the chosen voice — if is omitted

  3. 3

    Your app

    Receive the audio stream and play or save it

Basic Usage

from openai import OpenAI

client = OpenAI(
    api_key="cm_your_key",
    base_url="https://api.callmissed.com/v1"
)

response = client.audio.speech.create(
    model="bulbul:v3",
    voice="shubh",
    input="Namaste, kaise hain aap?"
)

response.stream_to_file("speech.mp3")

Parameters

Send a JSON body.

ParameterTypeRequiredDescription
inputstringYesText to synthesize, up to 4096 characters. Some models accept less: bulbul:v3 2500, bulbul:v2 1500, deepgram-aura-2 / deepgram-aura-1 2000
modelstringNoDefault bulbul:v3. One of bulbul:v3, sonic-3.6, gnani-timbre-v2.0, deepgram-aura-2, deepgram-aura-1, aura-2-en, aura-2-es, gpt-4o-mini-tts, melotts — see Models
voicestringNoVoice ID — default shubh for bulbul:v3, skylar for sonic-3.6 (see Voices). An unrecognized voice falls back to the model's default, except on gpt-4o-mini-tts, which returns 400 and lists its 13 voices
languagestringNoLanguage code (e.g. hi-IN, ta-IN; sonic-3.6 takes base codes like en, hi). bulbul:v3 defaults to en-IN
speednumberNoSpeech speed, default 1.0. Must be 0.25–4.0, then each model clamps to its own range: bulbul:v3 0.5–2.0, gpt-4o-mini-tts 0.25–4.0, deepgram-aura-2/-1 0.7–1.5, sonic-3.6 0.6–1.5, gnani-timbre-v2.0 0.85–1.15. Ignored by aura-2-en, aura-2-es, melotts
speech_sample_rateintegerNoDefault 8000. An integer from 8000 to 48000. bulbul:v3 accepts only 8000, 16000, 22050, 24000, 32000, 44100 or 48000 (anything else is a 400); Deepgram PCM output snaps to the nearest of 8000, 16000, 24000, 32000 or 48000
response_formatstringNoDefault mp3 — see Audio Formats
temperaturenumberNoExpressiveness, 0.01–2.0. bulbul:v3 only — higher is more expressive, lower is more consistent. Defaults to 0.9 (warmer than the model's flat default)
instructionsstringNoNatural-language delivery direction — tone, emotion, accent, pacing. gpt-4o-mini-tts only. Max 2000 chars. Example: "Speak slowly and warmly, like you're reassuring someone."
humanizebooleanNoDefault true. Shapes your text for natural speech before synthesis — strips markdown and emoji, speaks URLs and emails as words, groups long digit runs into readable chunks. Set false to synthesize your text byte-for-byte
streambooleanNodeepgram-aura-2 / deepgram-aura-1 only — stream audio as it's generated (lower time-to-first-byte)

Audio Formats

response_format accepts mp3, opus, aac, flac, wav or pcm, but not every model can produce every format. The Content-Type response header always describes the audio actually returned, so read it rather than assuming the format you asked for.

ModelFormats honouredOtherwise returns
bulbul:v3—Always WAV
aura-2-en, aura-2-es, melotts—Always MP3
gpt-4o-mini-ttsmp3, opus, aac, flac, wav, pcmMP3
deepgram-aura-2, deepgram-aura-1mp3, opus, aac, flac, wav, pcm, plus mulaw / alawWAV
sonic-3.6mp3, wav, pcmWAV
gnani-timbre-v2.0mp3, wav, pcm, opus, plus mulawWAV

Errors

Errors use the OpenAI envelope: {"error": {"message", "type", "code"}}.

StatuscodeWhen
400—The model rejected a parameter (for example an unsupported speech_sample_rate on bulbul:v3)
402insufficient_creditsCredit balance is exhausted
403permission_deniedThe API key does not have the tts permission, or the model is not available on your plan
404model_not_foundmodel is not a known TTS model ID
422invalid_requestMissing or empty input, input over the model's character limit (see input above), or a field with the wrong type or out of range (speed, speech_sample_rate, temperature, instructions)
429quota_exceededPlan usage limit reached
502provider_error, upstream_timeout, upstream_unavailableThe model failed to synthesize. The response includes a request_id

Making speech sound human

Naturalness comes from three places, in order of impact:

1. The text you send. Every engine sounds more human when the input reads like speech rather than like a screen. humanize (on by default) handles the mechanical part — markdown, emoji, https://callmissed.com → "callmissed dot com", 2039123456 → 203.912.3456. Beyond that, write short sentences and use contractions; if an LLM generates your text, tell it that its output will be spoken aloud.

2. Expressiveness parameters, where the model supports them:

# bulbul:v3 — temperature is its expressiveness control
curl -X POST https://api.callmissed.com/v1/audio/speech \
  -H "Authorization: Bearer cm_your_key" \
  -H "Content-Type: application/json" \
  -d '{"model": "bulbul:v3", "input": "Bilkul, main abhi check karta hoon.", "voice": "shubh", "temperature": 1.1}' \
  --output speech.mp3

# gpt-4o-mini-tts — direct the performance in plain language
curl -X POST https://api.callmissed.com/v1/audio/speech \
  -H "Authorization: Bearer cm_your_key" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4o-mini-tts", "input": "Your order shipped this morning.", "voice": "nova", "instructions": "Cheerful and upbeat, like sharing good news with a friend."}' \
  --output speech.mp3

3. Pauses in the text. deepgram-aura-2, deepgram-aura-1, aura-2-en and aura-2-es read pause cues written into your text: ... gives a longer natural pause, a comma or period gives a short one, and um/uh render as natural hesitation.

{"model": "deepgram-aura-2", "input": "Let me pull that up... okay, found it."}

bulbul:v3 and gnani-timbre-v2.0 do not support pause markup or SSML — they would speak the dots aloud. Use sentence length and real punctuation for rhythm on those models.

Choosing an expressive voice

WantUse
Most natural conversational speechsonic-3.6 — 44 languages, native-quality Hindi, sub-90ms first audio
Direct the emotion in wordsgpt-4o-mini-tts with instructions
Indian languages, warm deliverybulbul:v3 with temperature 0.9–1.2
Pauses and hesitation in textdeepgram-aura-2 (90 voices, many tagged expressive/cheerful)
Lowest costmelotts — no expressive controls; rely on humanize

Streaming (Deepgram Aura)

For deepgram-aura-2 and deepgram-aura-1, set "stream": true to receive audio frames as they're synthesized, as a chunked HTTP response. Ideal for real-time playback where you want the first audio bytes as fast as possible.

Streaming supports raw encodings only — response_format must be linear16 (or pcm/wav), mulaw, or alaw. Compressed formats (mp3, opus, aac, flac) cannot be streamed; if you request one with stream:true, the full audio is returned in one buffered response instead.

curl -N -X POST https://api.callmissed.com/v1/audio/speech \
  -H "Authorization: Bearer cm_your_key" \
  -H "Content-Type: application/json" \
  -d '{"model": "deepgram-aura-2", "voice": "thalia", "input": "Streaming hello.", "response_format": "linear16", "stream": true}' \
  --output speech.raw

Billing is identical to the non-streaming path (per character). Other providers ignore stream and return the full audio in one response.