Skip to main content

Speech to Text

Transcribe audio to text across 45 STT models — Indic-first saaras, Cartesia Ink, and Deepgram Nova.

Overview

Transcribe an audio file into text with any of 45 speech-to-text models. The default, saaras:v3, covers 22 Indian languages plus English with automatic language detection — but the model field takes any file-transcription STT id, so you can pick per request.

Endpoint: POST /v1/audio/transcriptions

Picking a model

If you needUseWhy
Indian languages, or code-mixed Hinglishsaaras:v3 or saaras:v4Purpose-built for Indic phonetics; v4 adds 24 languages and serves all five modes
The widest language coverageink-whisper100 languages, and cheaper per hour than the Indic models
English call-centre audiodeepgram-nova-3, or nova-3 on the free tierStrong on accented and noisy telephony English
Turn detection built into the modelink-2 or deepgram-flux-general-enVoice sessions only — see the note below

Four models are on the free tier: saaras:v3, saaras:v4, nova-3 and whisper-large-v3-turbo.

Some models appear under two ids at different prices — nova-3 and deepgram-nova-3 are the same underlying model, but only nova-3 is on the free tier. If you are on the free plan, use the id listed above; picking the other spelling of the same model will bill you. Full per-model pricing: Models.

How transcription works

  1. 1

    Your app

    Upload an audio file (WAV/MP3) to

  2. 2

    CallMissed gateway

    Validate the key, resolve , detect language (or use ), apply

  3. 3

    Your chosen model

    Run speech recognition — defaults to if is omitted

  4. 4

    Your app

    Receive (plus in )

Tip: Leave model unset to get saaras:v3, and leave language unset to let it auto-detect. Set mode=translate to get English text out of any supported language in a single call.

Streaming-only models cannot transcribe files. ink-2 and the Deepgram Flux models do turn detection as part of the model, which only makes sense on a live stream. Sending one here returns an error naming the file-transcription alternative rather than silently substituting a different model — see Cartesia Ink models.

Basic Usage

from openai import OpenAI

client = OpenAI(
    api_key="cm_your_key",
    base_url="https://api.callmissed.com/v1"
)

with open("audio.wav", "rb") as f:
    response = client.audio.transcriptions.create(
        model="saaras:v3",
        file=f
    )

print(response.text)

Parameters

Send as multipart/form-data.

ParameterTypeRequiredDescription
filefileYesAudio file (WAV, MP3, etc.), up to 25 MB. An empty file returns 400. Some models accept less per request: saaras:v3 and saaras:v4 take at most 30 seconds of audio, gnani-prisma-v2.5 at most 60 seconds, and whisper, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize at most 25,000,000 bytes
modelstringNoDefault saaras:v3. Any file-transcription STT model ID from Models. ink-2 and the Flux models are not valid here — see Cartesia Ink models
languagestringNoLanguage code; auto-detected if omitted. For saaras:v3 / saaras:v4 use a locale such as hi-IN or ta-IN — a code the model does not support falls back to auto-detection
modestringNoOutput mode — see below. Default transcribe; any other value returns 422
response_formatstringNojson (default), text, verbose_json, or diarized_json. Any other value is answered as json
temperaturenumberNo0–2. Accepted for OpenAI SDK compatibility; currently not forwarded to the model
promptstringNoAccepted for OpenAI SDK compatibility; currently not forwarded to the model

Response

json (default):

{"text": "Namaste, aap kaise hain?"}

text returns the transcript as text/plain. verbose_json adds the billed audio duration in seconds. words is always an empty array — word-level timestamps are not returned. segments is empty except on gpt-4o-transcribe-diarize, where each segment carries id, speaker, start, end and text. diarized_json returns task, duration, text and those speaker segments.

{
  "task": "transcribe",
  "language": "hi-IN",
  "duration": 4.52,
  "text": "Namaste, aap kaise hain?",
  "segments": [],
  "words": []
}

language echoes the language you sent, or "unknown" when you let the model auto-detect.

Errors

Errors use the OpenAI envelope: {"error": {"message", "type", "code"}}.

StatuscodeWhen
400invalid_requestfile is empty; the audio or a parameter was rejected (format, size, language); or the model is streaming-only or not available for file transcription. The message says which
400audio_too_longThe audio is longer than the model accepts per request (see file above). Use a model without a duration limit, such as deepgram-nova-3
402insufficient_creditsCredit balance is exhausted
403permission_deniedThe API key does not have the stt permission, or the model is not available on your plan
404model_not_foundmodel is not a known STT model ID
422—A form field fails validation (for example an unknown mode, or temperature outside 0–2)
413file_too_largeThe file is larger than 25 MB, or than the model accepts (see file above)
429quota_exceededPlan usage limit reached
429rate_limit_exceededTranscription is rate limited right now. Retry with backoff
502service_unavailableTranscription is temporarily unavailable. Retry shortly. /v1/audio/speech uses the same code
502provider_error, upstream_timeout, upstream_unavailableThe model failed to transcribe the file. The response includes a request_id

/v1/audio/translations returns the same errors.

Output Modes

ModeDescription
transcribeStandard transcription (default)
translateTranscribe and translate to English
verbatimExact transcription including filler words
translitTransliteration to Latin script
codemixCode-mixed output (Indic + English)

mode is applied by saaras:v3 and saaras:v4. whisper-large-v3-turbo also honours translate. Every other model ignores mode and returns a plain transcript.

saaras:v4 serves all five modes on one model across 24 languages, and is free-tier like saaras:v3:

curl -X POST https://api.callmissed.com/v1/audio/transcriptions \
  -H "Authorization: Bearer cm_your_key" \
  -F file=@audio.wav \
  -F model=saaras:v4 \
  -F mode=codemix

Cartesia Ink models

Two Cartesia STT models, and they are not interchangeable — one transcribes files, the other only runs on a live voice session.

ModelPriceLanguagesFile transcriptionVoice sessions
ink-whisper$0.1875 / hr100 (incl. Hindi, Urdu, Tamil)YesYes
ink-2$0.5625 / hr5: en, fr, hi, ja, esNoYes

ink-whisper — the cheapest 100-language option

Cartesia's fastest and most affordable STT, with better accuracy than baseline Whisper. Its dynamic chunking cuts hallucination during pauses and silence, so audio with dead air transcribes cleanly instead of inventing text to fill gaps.

curl -X POST https://api.callmissed.com/v1/audio/transcriptions \
  -H "Authorization: Bearer cm_your_key" \
  -F file=@audio.wav \
  -F model=ink-whisper \
  -F language=hi

ink-2 — voice agents only

Cartesia's top-ranked STT for voice agents: 8% WER on AppTek's 14-accent call-centre benchmark, against 10% for Deepgram Flux and 12% for ElevenLabs. It also self-detects turns, so a voice agent needs no separate turn detector on top.

Two limits decide whether you can use it at all:

1. It cannot transcribe files. ink-2 is streaming-only. POSTing it to /v1/audio/transcriptions fails rather than quietly substituting a different model. The request is not transcribed, and the error message names the alternative:

{
  "error": {
    "message": "ink-2 is a streaming-only model and is not available for file transcription. Use ink-whisper here, or ink-2 on a voice session.",
    "type": "invalid_request_error",
    "code": "invalid_request",
    "request_id": "stt-…"
  }
}

The HTTP status is 400. deepgram-flux-general-en and deepgram-flux-general-multi fail the same way and point you to deepgram-nova-3.

Select it on a voice session or the Managed Voice Agent instead.

2. It covers five languages. The model accepts en, fr, hi, ja and es. Sending it any other language (Tamil, for example) does not raise an upstream error — it silently produces poor output. Our agent logs a warning and transcribes as en. For other languages use ink-whisper (100 languages) or an Indic model such as saaras:v3 / saaras:v4.

Deepgram feature parameters

When you select a Deepgram file-transcription model (deepgram-nova-3, deepgram-nova-2, deepgram-enhanced, etc.), these extra form fields are accepted. They are ignored for non-Deepgram models. Model-restricted features are dropped automatically when the chosen model doesn't support them.

ParameterTypeDescription
diarizebooleanLabel each speaker ([Speaker 0], [Speaker 1], …)
utterancesbooleanSegment the transcript into utterances
utt_splitnumberSilence gap (seconds, 0–10) used to split utterances
paragraphsbooleanSplit the transcript into paragraphs
numeralsbooleanWrite numbers as digits (e.g. "five" → "5")
measurementsbooleanAbbreviate measurement units (English)
dictationbooleanConvert spoken "comma"/"period" to punctuation (English)
profanity_filterbooleanMask recognized profanity with ****
filler_wordsbooleanKeep "uh"/"um" (Nova / Nova-2 / Nova-3)
multichannelbooleanTranscribe each audio channel independently
detect_entitiesbooleanTag entities like names and locations (English)
detect_languagestringtrue to auto-detect, or repeat with codes to restrict
redactstringpci, pii, phi, numbers, or a specific entity type (repeatable)
keytermstringBoost recognition of a term/phrase (Nova-3; repeatable)
keywordsstringkeyword:intensifier boost/suppress (Nova-2 / legacy; repeatable)
searchstringPhonetically search the audio for a term (repeatable)
replacestringfind:replacement substitution (repeatable)

Dialects & locales

Deepgram models accept locale-specific language codes so you can pin a dialect for best accuracy. Pass the code in the language field. Each model's exact dialect list is published in the dialects array on GET /v1/models. Examples:

  • English: en-US, en-GB, en-IN, en-AU, en-NZ, en-CA, en-IE
  • Spanish: es, es-419 (Latin America)
  • Portuguese: pt-BR, pt-PT
  • Chinese: zh-CN, zh-TW, zh-HK (Cantonese)
  • Multilingual code-switching: multi (Nova-3, Nova-2)