Skip to main content

Models

Every model CallMissed serves — Indic STT/TTS/LLM, fast direct-routed LLMs, first-party flagships, realtime voice and image — through one OpenAI-compatible API, plus 300+ more we deploy on demand.

Overview

138 models, one OpenAI-compatible API. Same auth, same request shape — change the model field and nothing else.

GroupWhat it is
Fast LLMsGemma 4 31B, the default for voice agents, and Kimi K2.5 at up to ~414 tok/s.
Indic modelsSTT, TTS and LLM built for 22 Indian languages.
Direct-routed LLMsSub-2s open-weights models: Kimi K2.5/K2.6/K2.7 Code, GPT-OSS, Gemma 4, GLM, Nemotron, Mistral Small.
First-partygpt-4o, gpt-4.1, gpt-5-mini, gpt-5.5, gpt-5.6-*, grok-4.3, DeepSeek-V4-*, realtime voice, plus first-party STT/TTS.
On demand300+ more we deploy on request.

Models API

List all available models programmatically. No authentication required.

# List all models
curl https://api.callmissed.com/api/v1/models

# Filter by category: llm, stt, tts, image, embedding
curl https://api.callmissed.com/api/v1/models?category=llm

# Filter free-plan models only
curl https://api.callmissed.com/api/v1/models?free=true

# Get a specific model
curl https://api.callmissed.com/api/v1/models/sarvam-105b

# Which models each plan tier can call
curl https://api.callmissed.com/api/v1/models/access

Each entry includes id, name, description, category, owned_by, context_window, max_output_tokens (only where the model documents an output limit), pricing (US$), free, supports_streaming, supports_tools, supports_reasoning, supports_vision, zero_data_retention, hosted_in_india, voice_capable (valid for POST /v1/voice/sessions), voice_agent_only, and speed_tier. A model that is temporarily unavailable also carries status: "maintenance" and a status_message naming the alternative. Embedding models add dimensions; TTS models add their voice list. The response is {object: "list", total, data}.

The OpenAI-compatible listing at GET /v1/models (requires Authorization: Bearer cm_*) returns the same capability fields plus context_length (an alias of context_window for OpenAI-style clients), but a shorter list: it hides models that are only valid for voice sessions (nova-sonic*, gpt-realtime*, deepgram-voice-*). Use GET /api/v1/models for the full catalog. The Anthropic-shape listing at GET /anthropic/v1/models returns the same set inside Anthropic's {data, has_more, first_id, last_id} envelope.

Free Plan Models

The free tier includes 27 models across five categories. Use GET /api/v1/models?free=true to list them, or see the Model Access by Plan page for the full breakdown.

LLM (11 models)

Model IDDescription
sarvam-105b105B MoE — complex reasoning, Indic languages
sarvam-105b-conversations105B MoE tuned for conversation and voice — 32K context, tool calling
kimi-k2.5Moonshot K2.5 — 256K context, reasoning
kimi-k2.6Moonshot K2.6 — improved reasoning + coding, 262K context
kimi-k2.7-codeMoonshot K2.7 Code — frontier 1T-param agentic coding, 262K context, vision + tools
glm-4.7-flashGLM 4.7 Flash — fast inference
glm-5.2GLM 5.2 — Z.ai flagship agentic coding, 262K context, tools + reasoning
gpt-oss-120bGPT-OSS 120B — open-weights large model
nemotron-3-superNvidia Nemotron 3 Super
gemma-4-26b-a4b-itGoogle Gemma 4 26B
mistral-small-3.1Mistral Small 3.1 — 24B instruct, tool use

STT (4 models)

Model IDDescription
saaras:v323 langs (22 Indic + English), best for code-mixed
saaras:v424 langs — five output modes: transcribe, translate, verbatim, transliterate, code-mix
whisper-large-v3-turboWhisper — 99 langs with auto-detect, transcribe + translate
nova-3Nova 3 — 60+ languages, diarization, smart-format, streaming-capable

TTS (4 models)

Model IDDescription
bulbul:v337 voices, 11 Indian languages
aura-2-enAura 2 — 40 English voices, low-latency streaming
aura-2-esAura 2 — 10 Spanish voices, low-latency streaming
melottsMeloTTS — en + fr, cheapest TTS available

Image (6 models)

Model IDDescription
flux-2-klein-9bFlux 2 Klein — highest quality
flux-2-devFlux 2 Dev — flagship fidelity
lucid-originLucid Origin — cinematic
phoenix-1.0Phoenix — photorealistic
sdxl-lightningSDXL Lightning — fast
dreamshaper-8-lcmDreamShaper 8 LCM — fast

Embedding (2 models)

Model IDDescription
text-embedding-3-small1536 dimensions, 8,192-token inputs — best price/performance
text-embedding-3-large3072 dimensions, 8,192-token inputs — highest accuracy

A free-plan key calling a paid model gets 403 model_not_available — it is not billed, it is refused. Upgrade to Starter or above first.

All other models — including kimi-k2.5-fast, glm-5.3, gemma-4-31b, the gemini-* chat models, first-party IDs (gpt-4o, gpt-4.1, gpt-5-mini, gpt-5.5, gpt-5.6-*, gpt-6-*, gpt-6.1-sol, grok-4.3, DeepSeek-V4-*, gpt-realtime*, gpt-live-1, nova-sonic*, first-party STT/TTS), the other speech models (gnani-*, ink-*, sonic-3.6), the Deepgram direct line (deepgram-nova-3, deepgram-flux-general-en/multi, deepgram-nova-2*, deepgram-enhanced*, deepgram-base*, deepgram-whisper-*, deepgram-aura-2, deepgram-aura-1, Deepgram Voice Agent deepgram-voice-* ids, the deepgram-summarize/topics/sentiment/intents Audio Intelligence features, and the deepgram-text-summarize/topics/sentiment/intents Text Intelligence features), and paid image models (flux-2-pro, flux-1.1-pro, gpt-image-2.5-*, gpt-image-2, gpt-image-1.5, nano-banana-*, gemini-3.1-flash-lite-image) — require Starter, Pro, or Enterprise.

Pricing

All models are pay-per-use. Pricing is in US$ at US$1 = ₹96; you pay in credits, where 1 credit = ₹1 ≈ US$0.0104.

ModelInput / 1M tokensOutput / 1M tokens
kimi-k2.5-fast$0.8438$4.219
sarvam-105b$0.3646 (₹35)$0.3646 (₹35)
sarvam-105b-conversations$0.3646 (₹35)$0.3646 (₹35)
gpt-5.6-sol$5.208$31.25
gpt-5.6-terra$2.083$12.5
gpt-5.6-luna$0.2083$1.25
gpt-6-sol$2.083$10.42
gpt-6-luna$0.1042$0.5208
gpt-6.1-sol$2.083$10.42
nova-sonic-2$4.167$15.63
nova-sonic$4.688$17.71
gpt-realtime$4.167$16.67
gpt-realtime-mini$0.625$2.5
gpt-realtime-2$4.167$25
gpt-realtime-1.5$4.167$16.67
gpt-realtime-2.1$4.167$25
gpt-realtime-2.1-mini$0.625$2.5
gpt-live-1per-minute only$0.05208/min (voice-agent only)
deepgram-voice-*per-minute Voice Agent tierStandard $0.07813/min, Advanced $0.1698/min (voice-agent only)
STT ModelPrice
saaras:v3$0.3125 / hour (₹30/hr)
saaras:v4$0.3125 / hour (₹30/hr)
gnani-prisma-v2.5$0.2813 / hour
ink-whisper$0.1875 / hour
ink-2$0.5625 / hour (voice sessions only)
whisper-large-v3-turbo$0.0625 / hour
nova-3$0.5208 / hour
deepgram-nova-3$0.3021 / hour
deepgram-nova-3-medical$0.3021 / hour
deepgram-flux-general-en$0.4063 / hour
deepgram-flux-general-multi$0.4896 / hour
deepgram-nova-2 (+ domain variants)$0.3646 / hour
deepgram-nova / deepgram-whisper-*$0.3646 / hour
deepgram-enhanced (+ variants)$1.031 / hour
deepgram-base (+ variants)$0.9063 / hour
TTS ModelPrice
bulbul:v3$0.3125 / 10K chars (₹30/10K)
gnani-timbre-v2.0$0.2813 / 10K chars
aura-2-en$0.4167 / 10K chars
aura-2-es$0.4167 / 10K chars
deepgram-aura-2$0.3125 / 10K chars
deepgram-aura-1$0.1563 / 10K chars
sonic-3.6$0.5208 / 10K chars
melotts$0.05208 / 10K chars
Intelligence feature (not a model ID)Price
deepgram-summarize / -topics / -sentiment / -intents (audio)$0.0003125 / 1K input + $0.000625 / 1K output tokens
deepgram-text-summarize / -text-topics / -text-sentiment / -text-intents$0.0003125 / 1K input + $0.000625 / 1K output tokens

These eight values go in the features field, not in model. They are not catalog models — GET /api/v1/models/deepgram-summarize returns 404.

Full pricing for all models is available via the API: GET /api/v1/models

import requests

# List all LLM models
models = requests.get("https://api.callmissed.com/api/v1/models?category=llm").json()
for m in models["data"]:
    print(f"{m['id']} — {m['name']} ({m['context_window']} tokens) {'FREE' if m['free'] else 'PAID'}")

Fast LLMs

High-throughput Kimi K2.5 inference tier optimized for voice-agent latency.

Model IDStatusContextBest For
kimi-k2.5-fastUnder maintenance — fall back to kimi-k2.5256KVoice agents, fast inference, reasoning tasks

While kimi-k2.5-fast is in maintenance (returns HTTP 503), use kimi-k2.5:

response = client.chat.completions.create(
    model="kimi-k2.5",
    messages=[{"role": "user", "content": "Hello"}]
)

Indic Models

Speech to Text

ModelDescriptionLanguages
saaras:v3Latest STT — best accuracy on Indian + code-mixed23 languages (22 Indic + English)
saaras:v4Five output modes on one model — transcribe, translate, verbatim, transliterate, code-mix24 languages
gnani-prisma-v2.5India-first telephony STT — code-switching, sub-4% WER on Indian English10 Indian languages

For 99-language general-purpose transcription, see whisper-large-v3-turbo. For diarization + smart-format on calls, see nova-3. Both are free-tier and live under the audio model routes.

ink-whisper also covers Hindi, Urdu and Tamil as part of its 100-language set at $0.1875 / hr — cheaper than the Indic-specialist models, though without their code-mix output modes. ink-2 covers Hindi among its five languages (en, fr, hi, ja, es) but no other Indic language.

Text to Speech

ModelDescriptionVoices
bulbul:v3Natural TTS — 37 voices, 11 Indian languagesshubh (default) + 36 more
gnani-timbre-v2.0India-first neural TTS — context-aware tone, low-latency. Served by Gnani's current Timbre generation (v2.5); the id keeps its original name73 voices (English + Hindi + Indic)
sonic-3.6Cartesia Sonic 3.6 — most natural conversational speech, 44 languages with native-quality HindiSearchable public library + 16 featured aliases (skylar default)

For low-latency English / Spanish voice agents, see aura-2-en / aura-2-es. For ultra-cheap en/fr notification audio, see melotts. All three are free-tier.

Chat Completion (LLM)

ModelParamsContextBest For
sarvam-105b105B MoE128K tokensComplex reasoning, agentic tasks, long documents
sarvam-105b-conversations105B MoE32K tokensConversation and voice agents, tool calling
glm-5.3Z.ai GLM 5.31M tokensLong-context reasoning, tool calling (text only)
gemma-4-31bGoogle Gemma 4 31B128K tokensFast answers with thinking off, image input, tool calling, voice agents

glm-5.3 always reasons. reasoning_effort takes "low", "high" or "max" (the default when omitted); "none" and "minimal" map to "low", "medium" to "high" and "xhigh" to "max". Reasoning tokens bill as output. gemma-4-31b answers without thinking by default. reasoning_effort: "low", "medium" or "high" switch thinking on; the level does not reliably change how much it thinks. "none" and "minimal" keep it off. Both are paid-plan models.

Both sarvam-105b models support hybrid thinking via reasoning_effort: "low" | "high" | "max", and "xhigh" maps to "max". "none" and "minimal" switch thinking fully off. "medium" is not one of their levels, so it is dropped and the model runs at its default. Other models with a full off switch: kimi-k2.5, kimi-k2.6, glm-4.7-flash, glm-5.2, gemma-4-26b-a4b-it, gemma-4-31b, DeepSeek-V4-Pro, DeepSeek-V4-Flash, or a GPT-5.5 / GPT-5.6 / GPT-6 model with reasoning_effort: "none". See the per-model matrix.

Audio Models

Free-tier on every plan. See the Pricing page for current rates.

Speech to Text

ModelLanguagesBest forPrice
whisper-large-v3-turbo99 with auto-detectMultilingual general-purpose; transcribe + translate$0.0625 / hour
nova-311 BCP-47 incl. multi auto-detectDiarization, smart-format, streaming voice agents$0.5208 / hour
whisper99 with auto-detectWhisper batch + translate$0.4167 / hour
gpt-4o-transcribeStreamingHigher-accuracy OpenAI transcription$0.4167 / hour
gpt-4o-mini-transcribeStreamingLow-cost OpenAI transcription$0.25 / hour
gpt-4o-transcribe-diarizeFile transcription with speaker labels (not available for voice sessions)Multi-speaker meetings / calls$0.4167 / hour

Deepgram (direct) — the full Deepgram speech-to-text line, billed per audio hour at the rates below:

ModelLanguagesBest forPrice
deepgram-flux-general-enEnglishConversational voice agents — model-native turn detection, ultra-low latency$0.4063 / hour
deepgram-flux-general-multi10 (multilingual)Multilingual voice agents with code-switching$0.4896 / hour
deepgram-nova-345+ incl. multiFlagship general-purpose ASR, keyterm prompting, PII redaction$0.3021 / hour
deepgram-nova-3-medicalEnglishClinical / medical terminology$0.3021 / hour
deepgram-nova-236 incl. multiHigh-accuracy ASR + filler-word detection$0.3646 / hour
deepgram-nova-2-{meeting,phonecall,finance,conversationalai,voicemail,video,medical,drivethru,automotive,atc}EnglishDomain-tuned Nova-2 variants$0.3646 / hour
deepgram-nova / -phonecall / -medicalen/es/hiLegacy Nova-1$0.3646 / hour
deepgram-enhanced (+ meeting/phonecall/finance)13Legacy, keyword boosting$1.031 / hour
deepgram-base (+ 6 variants)17Legacy, high-volume batch$0.9063 / hour
deepgram-whisper-{tiny,base,small,medium,large}99Whisper in five sizes, batch transcription$0.3646 / hour

Text to Speech

ModelLanguagesVoicesPrice
aura-2-enEnglish40 (luna default)$0.4167 / 10K chars
aura-2-esSpanish10 (aquila default)$0.4167 / 10K chars
deepgram-aura-2en/es/de/fr/nl/it/ja90+ (thalia default)$0.3125 / 10K chars
deepgram-aura-1English12 (asteria default)$0.1563 / 10K chars
melottsEnglish + French1 per language$0.05208 / 10K chars
gpt-4o-mini-ttsMultilingual steerable6 OpenAI voices$0.2083 / 10K chars

Aura 2 returns linear16 PCM streamed at 24 kHz for low-latency playback. MeloTTS returns base64 MP3. Output formats may vary as models are updated.

Deepgram Flux TTS is a voice-agent-first model and is not available on this /v1/audio/speech endpoint. It is offered only through the managed Voice Agent (see the Voice Sessions API), selectable with tts_engine: "flux", where it is billed inside the per-minute voice rate.

Audio Intelligence (Deepgram)

Deepgram Audio Intelligence runs analysis over an uploaded audio file via POST /v1/audio/intelligence (English only, 150K input-token limit). Token-billed at $0.0003125/1K input + $0.000625/1K output.

FeatureModel IDReturns
Summarizationdeepgram-summarizeA concise summary of the audio
Topic Detectiondeepgram-topicsPer-segment topics with confidence
Sentiment Analysisdeepgram-sentimentPer-segment + average sentiments
Intent Recognitiondeepgram-intentsPer-segment intents with confidence

Request multiple features in one call with a comma-separated features form field (e.g. features=deepgram-summarize,deepgram-sentiment).

Text Intelligence (Deepgram)

Deepgram Text Intelligence runs the same four analyses over text input (a string or a hosted text URL) via POST /v1/text/intelligence (English only, 150K input-token limit). Token-billed at $0.0003125/1K input + $0.000625/1K output. Requires the llm key permission.

FeatureModel IDReturns
Summarizationdeepgram-text-summarizeA concise summary of the text
Topic Detectiondeepgram-text-topicsPer-segment topics with confidence
Sentiment Analysisdeepgram-text-sentimentPer-segment + average sentiments
Intent Recognitiondeepgram-text-intentsPer-segment intents with confidence

Send a JSON body with features (array or comma-separated string) and exactly one of text or url:

{
  "features": ["deepgram-text-summarize", "deepgram-text-sentiment"],
  "text": "Your text to analyze here."
}

Direct-Routed LLMs

Low-latency models routed directly through CallMissed — sub-2s end-to-end on small prompts and free-tier eligible per the reasoning_effort matrix.

Model IDCreatorContext
kimi-k2.5Moonshot AI256K
kimi-k2.6Moonshot AI262K
kimi-k2.7-codeMoonshot AI262K
gpt-oss-120bOpenAI (open-weights)128K
gemma-4-26b-a4b-itGoogle256K
glm-4.7-flashZhipu128K
glm-5.2Z.ai262K
nemotron-3-superNVIDIA256K
mistral-small-3.1Mistral128K

Models on Demand

GET /api/v1/models lists everything in the catalogue today: 138 model IDs. Models marked status: "maintenance" return 503 until they are back.

Beyond that we deploy 300+ further models on demand on CallMissed infrastructure. Send the model you need and your expected throughput to sales@callmissed.com. Once deployed it appears in your GET /api/v1/models response with a plain CallMissed ID and published per-token pricing, on the same /v1/chat/completions endpoint as every other model. Same key, same credit balance, no new SDK.

Enterprise accounts get dedicated capacity. Starter and Pro get shared capacity where the model allows it.

First-Party Models

Credit-covered first-party models. Use the bare model ID in API requests — e.g. gpt-4o.

Model IDTypeNotes
gpt-4oLLMMultimodal text + vision, 128K context
gpt-4.1LLMLong-context (300K) multimodal
gpt-5-miniLLMFast reasoning, 400K context
gpt-5.5LLMGPT-5.5 reasoning flagship, 1.05M context, vision + tools
gpt-5.6-solLLMGPT-5.6 flagship, 1.05M context, vision + tools
gpt-5.6-terraLLMGPT-5.6 balanced intelligence/cost, 1.05M context
gpt-5.6-lunaLLMGPT-5.6 fast + affordable, 1.05M context
gpt-6-solLLMGPT-6 frontier reasoning, 1.05M context, vision + tools
gpt-6-lunaLLMGPT-6 efficient, high-volume, 1.05M context
gpt-6.1-solLLMGPT-6.1 near-frontier for complex coding and professional work, 1.05M context, vision + tools
grok-4.3LLMxAI Grok, 200K context
DeepSeek-V4-ProLLMFlagship DeepSeek reasoning, 1M context, tools (text-only)
DeepSeek-V4-FlashLLMFast DeepSeek reasoning, 1M context, tools
nova-sonic-2Realtime voiceAmazon Nova 2 speech-to-speech voice model — 16 voices, Hindi + en-IN, live
nova-sonicRealtime voiceFirst-generation Amazon speech-to-speech voice model — retired upstream 2026-09-14, use nova-sonic-2
gpt-realtimeRealtime voiceOpenAI flagship speech-to-speech (10 concurrent), live
gpt-realtime-miniRealtime voiceLowest-cost realtime, ~3× cheaper than gpt-realtime, live
gpt-realtime-2Realtime voiceNewest realtime with stronger tool calling, live
gpt-realtime-1.5Realtime voicePinned 1.5 snapshot of gpt-realtime, live
gpt-realtime-2.1Realtime voiceLatest realtime — better recognition, silence/interrupt handling, live
gpt-realtime-2.1-miniRealtime voiceDistilled low-cost 2.1 realtime, live
gpt-live-1Realtime voiceGPT Live speech-to-speech — audio + text only, 14 voices, $0.05208/min, live
whisperSTTOpenAI Whisper — 99 langs
gpt-4o-transcribeSTTStreaming transcription
gpt-4o-mini-transcribeSTTLow-cost streaming STT
gpt-4o-transcribe-diarizeSTTSpeaker diarization
gpt-4o-mini-ttsTTSSteerable OpenAI TTS, 13 voices

See Credits & Rate Limits for per-model USD pricing.

Full Model Catalog

A curated, representative slice of the 138 models (67 LLM · 45 STT · 9 TTS · 15 image · 2 embedding) served by GET /api/v1/models as of the latest deploy — the per-domain variants of the direct Deepgram STT line and the deepgram-voice-* managed LLM ids are covered in their own sections above rather than repeated below. For live pricing and capability flags (supports_vision, supports_tools, free), query the API — it always reflects the current catalog.

LLM (43 models)

Model IDDescriptionContextFreePricing
sarvam-105b105B MoE. Complex reasoning, agentic tasks, long documents.131KYes$0.3646 in / $0.3646 out per 1M
sarvam-105b-conversations105B MoE tuned for conversation and voice. Tool calling.32KYes$0.3646 in / $0.3646 out per 1M
gpt-4oMultimodal text + vision.128KNo$2.604 in / $10.42 out per 1M
gemini-3.8-flashFast multimodal flagship. Thinking low/medium/high.1MNo$1.563 in / $7.813 out per 1M
gemini-3.7-flashFast multimodal. Thinking low/medium/high.1MNo$1.563 in / $7.813 out per 1M
gemini-3.6-flashFast multimodal. Thinking minimal→high.1MNo$1.563 in / $7.813 out per 1M
gemini-3.5-flashBalanced multimodal workhorse.1MNo$1.563 in / $9.375 out per 1M
gemini-3.5-flash-liteCheapest 1M-context Gemini.1MNo$0.3125 in / $2.604 out per 1M
gemini-3.1-pro-previewReasoning-heavy Gemini tier.1MNo$2.083 in / $12.5 out per 1M
gemini-3.1-flash-liteLow-cost multimodal, tool use.1MNo$0.2604 in / $1.563 out per 1M
gpt-4.1Long-context multimodal. Strong instruction following.300KNo$2.083 in / $8.333 out per 1M
gpt-5-miniFast, affordable reasoning.400KNo$0.2604 in / $2.083 out per 1M
gpt-5.5Reasoning flagship. Vision, tools, prompt caching.1.05MNo$5.208 in / $31.25 out per 1M
gpt-5.6-solFrontier model for complex professional work. Vision, reasoning, tools.1.05MNo$5.208 in / $31.25 out per 1M
gpt-5.6-terraBalances intelligence and cost. Vision, reasoning, tools.1.05MNo$2.083 in / $12.5 out per 1M
gpt-5.6-lunaCost-sensitive, high-volume workloads. Vision, reasoning, tools.1.05MNo$0.2083 in / $1.25 out per 1M
gpt-6-solFrontier reasoning for enterprise agents, coding and complex knowledge work. Vision, reasoning, tools.1.05MNo$2.083 in / $10.42 out per 1M
gpt-6-lunaEfficient GPT-6 for high-volume, cost-sensitive workloads. Vision, reasoning, tools.1.05MNo$0.1042 in / $0.5208 out per 1M
gpt-6.1-solNear-Astra performance for complex coding, computer use and professional work at a lower cost. Vision, reasoning, tools.1.05MNo$2.083 in / $10.42 out per 1M
grok-4.3xAI Grok 4.3. Reasoning + vision (images of at least 512 pixels in total; smaller returns 400). Reasoning tokens bill as output.200KNo$3.646 in / $15.63 out per 1M
DeepSeek-V4-ProFlagship DeepSeek reasoning. Tools. Text-only.1MNo$1.375 in / $4.125 out per 1M
DeepSeek-V4-FlashFast, affordable DeepSeek reasoning. Tools.1MNo$0.4583 in / $1.375 out per 1M
kimi-k2.5Strong on coding and math. Vision.256KYes$0.8438 in / $4.219 out per 1M
kimi-k2.5-fast (maintenance)Kimi K2.5 at ~414 tok/s for voice-agent latency.256KNo$0.8438 in / $4.219 out per 1M
kimi-k2.6Improved reasoning and coding over K2.5. Vision.262KYes$1.333 in / $5.625 out per 1M
kimi-k2.7-code1T-param agentic coding. Vision + tools.262KYes$1.333 in / $5.625 out per 1M
glm-4.7-flashFast, cost-efficient bilingual model. Strong tool use.131KYes$0.5208 in / $2.083 out per 1M
glm-5.2Flagship agentic coding. Tools + reasoning.262KYes$1.969 in / $6.188 out per 1M
glm-5.3Z.ai GLM 5.3. Tools + reasoning (low / high / max). Text-only.1MNo$2.125 in / $6.667 out per 1M
gpt-oss-120bOpen-weight 120B MoE. Reasoning-grade at lower cost.128KYes$1.042 in / $4.167 out per 1M
nemotron-3-super120B MoE tuned for long-context reasoning.256KYes$1.563 in / $6.25 out per 1M
gemma-4-26b-a4b-it26B MoE (4B active). Efficient instruct model. Vision.256KYes$0.4167 in / $1.667 out per 1M
gemma-4-31bGemma 4 31B instruct. Fast with thinking off, vision, tools.128KNo$0.6146 in / $1.542 out per 1M
mistral-small-3.124B instruct. Strong tool use, fast. Vision.128KYes$0.4896 in / $0.7917 out per 1M
nova-sonic-2Amazon Nova 2 Sonic. Native speech-to-speech voice model — STT, reasoning, and TTS in one; 16 voices across 8 languages including Hindi + en-IN.1MNo$4.167 in / $15.63 out per 1M • $0.06667/min
nova-sonicAmazon Nova Sonic 1.0. Native speech-to-speech voice model with 11 voices across English, Spanish, French, Italian, and German.32KNo$4.688 in / $17.71 out per 1M • $0.07396/min
gpt-realtimeOpenAI flagship realtime speech-to-speech model — STT + reasoning + function calling + TTS in one. 10 concurrent.32KNo$4.167 in / $16.67 out per 1M • $0.3906/min
gpt-realtime-miniLowest-cost realtime — same single-model shape as gpt-realtime, ~3× cheaper. 20 concurrent.32KNo$0.625 in / $2.5 out per 1M • $0.1229/min
gpt-realtime-2Newest realtime with stronger tool calling. 32K input, 4K output.32KNo$4.167 in / $25 out per 1M • $0.3906/min
gpt-realtime-1.5Pinned 1.5 snapshot of gpt-realtime. Use when you want version stability.32KNo$4.167 in / $16.67 out per 1M • $0.3906/min
gpt-realtime-2.1Latest realtime speech-to-speech — better alphanumeric recognition, silence/noise + interruption handling. Voice-agent only.32KNo$4.167 in / $25 out per 1M • $0.3906/min
gpt-realtime-2.1-miniDistilled, lower-cost realtime for faster voice interactions. Voice-agent only.32KNo$0.625 in / $2.5 out per 1M • $0.1219/min
gpt-live-1GPT Live speech-to-speech — audio + text in and out, function calling, 14 voices (default marin). No image or video input. Voice-agent only.—No$0.05208/min (billed per second)

Speech to Text (11 models)

Model IDDescriptionContextFreePricing
saaras:v323 languages (22 Indic + English). Best on code-mixed speech.—Yes$0.3125 / hr
saaras:v424 languages. Five output modes: transcribe, translate, verbatim, transliterate, code-mix.—Yes$0.3125 / hr
gnani-prisma-v2.5India-first telephony STT. 10 Indian languages, code-switching.—No$0.2813 / hr
ink-whisperCartesia Ink Whisper — 100 languages including Hindi, Urdu and Tamil. Dynamic chunking reduces hallucination across pauses and silence. File transcription + streaming.—No$0.1875 / hr
ink-2Cartesia Ink 2 — top-ranked for voice agents (8% WER on AppTek's 14-accent call-centre benchmark, vs 10% Deepgram Flux and 12% ElevenLabs). Self-detects turns. Languages: en, fr, hi, ja, es. Voice sessions only — not available for file transcription.—No$0.5625 / hr
whisper-large-v3-turbo99 languages with auto-detect. Transcribe + translate.—Yes$0.0625 / hr
nova-3Diarization, punctuation, smart-format. Streaming-capable.—Yes$0.5208 / hr
whisper99 languages. Transcription + translation to English.—No$0.4167 / hr
gpt-4o-transcribeHigher accuracy than Whisper. Streaming.—No$0.4167 / hr
gpt-4o-mini-transcribeCheaper, faster streaming transcription.—No$0.25 / hr
gpt-4o-transcribe-diarizeFile transcription with speaker labels. Not available for voice sessions.—No$0.4167 / hr

Text to Speech (7 models)

Model IDDescriptionVoicesFreePricing
bulbul:v3Indic TTS across 11 Indian languages.37Yes$0.3125 / 10K chars
gnani-timbre-v2.0India-first neural TTS, English + Hindi + Indic. Context-aware tone. Served by Timbre v2.5.73No$0.2813 / 10K chars
sonic-3.6Cartesia Sonic 3.6 — most natural conversational TTS. 44 languages, native-quality Hindi + Hinglish, sub-90ms first audio.Searchable libraryNo$0.5208 / 10K chars
aura-2-enConversational English TTS, low-latency streaming.40Yes$0.4167 / 10K chars
aura-2-esSpanish TTS, low-latency streaming.10Yes$0.4167 / 10K chars
melottsLightweight English + French TTS. Cheapest available.1 per languageYes$0.05208 / 10K chars
gpt-4o-mini-ttsSteerable — takes an instructions field to direct tone.6No$0.2083 / 10K chars

Image Generation (15 models)

Model IDDescriptionFreePricing
flux-2-klein-9bFlux 2 Klein. 1024×1024 default.Yes$0.1042 / image
flux-2-devFlux 2 Dev. Higher fidelity, 50-step inference.Yes$0.125 / image
flux-2-proFlux 2 Pro. Flagship BFL fidelity.No$0.1042 / image
flux-1.1-proFlux 1.1 Pro. Fast, production-grade.No$0.05208 / image
gpt-image-2.5-sunburstMost capable generation + editing. Inpainting, quality tiers to max.No$0.2604 / image
gemini-3.1-flash-lite-imageLow-cost text-to-image with reference edits. 1K resolution.No$0.035 / image
gpt-image-2.5-flareFast, high-quality everyday generation. On-image text + edits.No$0.2604 / image
gpt-image-2Accurate on-image text rendering.No$0.2604 / image
gpt-image-1.5Precise image editing. Strong logo/face preservation.No$0.2604 / image
lucid-originVibrant, cinematic compositions.Yes$0.08333 / image
phoenix-1.0Strong prompt adherence, photorealistic portraits.Yes$0.1042 / image
sdxl-lightning4-step inference. Fastest for iterative prompting.Yes$0.04167 / image
dreamshaper-8-lcmStylised illustrations, fast generation.Yes$0.04167 / image
nano-banana-2Fast multimodal image generation.No$0.06979 / image
nano-banana-proFlagship typography and fidelity.No$0.1396 / image

Embeddings (2 models)

Model IDDescriptionDimensionsFreePricing
text-embedding-3-smallFast, low-cost embeddings. Best price/performance for large corpora.1536Yes$0.02083 / 1M input tokens
text-embedding-3-largeHighest-accuracy embeddings.3072Yes$0.1354 / 1M input tokens

Both accept 8,192-token inputs and support shortening the vector with dimensions. See Embeddings.

Tip: Filter programmatically — GET /api/v1/models?category=llm, ?category=stt, ?category=tts, ?category=image, ?category=embedding, or ?free=true for free-plan models only.

Model Selection

Pass the model ID in your request:

# Indic LLM
response = client.chat.completions.create(
    model="sarvam-105b",
    messages=[{"role": "user", "content": "Hello in Hindi"}]
)

# Indic LLM with thinking mode
response = client.chat.completions.create(
    model="sarvam-105b",
    messages=[{"role": "user", "content": "Solve this step by step"}],
    extra_body={"reasoning_effort": "high"}
)

# First-party flagship model
response = client.chat.completions.create(
    model="gpt-5.6-luna",
    messages=[{"role": "user", "content": "Hello"}]
)

The API automatically routes to the correct backend based on the model ID:

  • Bare names (kimi-k2.5, gpt-4o, DeepSeek-V4-Pro, mistral-small-3.1, …) → direct-routed or first-party
  • sarvam-* prefix → Indic LLMs
  • saaras:* / bulbul:* / deepgram-* / image IDs → the matching speech or image backend

Every ID is a plain CallMissed ID with no vendor prefix — including models we deploy on demand.