Models
Every model CallMissed serves — Indic STT/TTS/LLM, fast direct-routed LLMs, first-party flagships, realtime voice and image — through one OpenAI-compatible API, plus 300+ more we deploy on demand.
Overview
138 models, one OpenAI-compatible API. Same auth, same request shape — change
the model field and nothing else.
| Group | What it is |
|---|---|
| Fast LLMs | Gemma 4 31B, the default for voice agents, and Kimi K2.5 at up to ~414 tok/s. |
| Indic models | STT, TTS and LLM built for 22 Indian languages. |
| Direct-routed LLMs | Sub-2s open-weights models: Kimi K2.5/K2.6/K2.7 Code, GPT-OSS, Gemma 4, GLM, Nemotron, Mistral Small. |
| First-party | gpt-4o, gpt-4.1, gpt-5-mini, gpt-5.5, gpt-5.6-*, grok-4.3, DeepSeek-V4-*, realtime voice, plus first-party STT/TTS. |
| On demand | 300+ more we deploy on request. |
Models API
List all available models programmatically. No authentication required.
# List all models
curl https://api.callmissed.com/api/v1/models
# Filter by category: llm, stt, tts, image, embedding
curl https://api.callmissed.com/api/v1/models?category=llm
# Filter free-plan models only
curl https://api.callmissed.com/api/v1/models?free=true
# Get a specific model
curl https://api.callmissed.com/api/v1/models/sarvam-105b
# Which models each plan tier can call
curl https://api.callmissed.com/api/v1/models/accessEach entry includes id, name, description, category, owned_by, context_window, max_output_tokens (only where the model documents an output limit), pricing (US$), free, supports_streaming, supports_tools, supports_reasoning, supports_vision, zero_data_retention, hosted_in_india, voice_capable (valid for POST /v1/voice/sessions), voice_agent_only, and speed_tier. A model that is temporarily unavailable also carries status: "maintenance" and a status_message naming the alternative. Embedding models add dimensions; TTS models add their voice list. The response is {object: "list", total, data}.
The OpenAI-compatible listing at GET /v1/models (requires Authorization: Bearer cm_*) returns the same capability fields plus context_length (an alias of context_window for OpenAI-style clients), but a shorter list: it hides models that are only valid for voice sessions (nova-sonic*, gpt-realtime*, deepgram-voice-*). Use GET /api/v1/models for the full catalog. The Anthropic-shape listing at GET /anthropic/v1/models returns the same set inside Anthropic's {data, has_more, first_id, last_id} envelope.
Free Plan Models
The free tier includes 27 models across five categories. Use GET /api/v1/models?free=true to list them, or see the Model Access by Plan page for the full breakdown.
LLM (11 models)
| Model ID | Description |
|---|---|
sarvam-105b | 105B MoE — complex reasoning, Indic languages |
sarvam-105b-conversations | 105B MoE tuned for conversation and voice — 32K context, tool calling |
kimi-k2.5 | Moonshot K2.5 — 256K context, reasoning |
kimi-k2.6 | Moonshot K2.6 — improved reasoning + coding, 262K context |
kimi-k2.7-code | Moonshot K2.7 Code — frontier 1T-param agentic coding, 262K context, vision + tools |
glm-4.7-flash | GLM 4.7 Flash — fast inference |
glm-5.2 | GLM 5.2 — Z.ai flagship agentic coding, 262K context, tools + reasoning |
gpt-oss-120b | GPT-OSS 120B — open-weights large model |
nemotron-3-super | Nvidia Nemotron 3 Super |
gemma-4-26b-a4b-it | Google Gemma 4 26B |
mistral-small-3.1 | Mistral Small 3.1 — 24B instruct, tool use |
STT (4 models)
| Model ID | Description |
|---|---|
saaras:v3 | 23 langs (22 Indic + English), best for code-mixed |
saaras:v4 | 24 langs — five output modes: transcribe, translate, verbatim, transliterate, code-mix |
whisper-large-v3-turbo | Whisper — 99 langs with auto-detect, transcribe + translate |
nova-3 | Nova 3 — 60+ languages, diarization, smart-format, streaming-capable |
TTS (4 models)
| Model ID | Description |
|---|---|
bulbul:v3 | 37 voices, 11 Indian languages |
aura-2-en | Aura 2 — 40 English voices, low-latency streaming |
aura-2-es | Aura 2 — 10 Spanish voices, low-latency streaming |
melotts | MeloTTS — en + fr, cheapest TTS available |
Image (6 models)
| Model ID | Description |
|---|---|
flux-2-klein-9b | Flux 2 Klein — highest quality |
flux-2-dev | Flux 2 Dev — flagship fidelity |
lucid-origin | Lucid Origin — cinematic |
phoenix-1.0 | Phoenix — photorealistic |
sdxl-lightning | SDXL Lightning — fast |
dreamshaper-8-lcm | DreamShaper 8 LCM — fast |
Embedding (2 models)
| Model ID | Description |
|---|---|
text-embedding-3-small | 1536 dimensions, 8,192-token inputs — best price/performance |
text-embedding-3-large | 3072 dimensions, 8,192-token inputs — highest accuracy |
A free-plan key calling a paid model gets 403 model_not_available — it is not billed, it is refused. Upgrade to Starter or above first.
All other models — including kimi-k2.5-fast, glm-5.3, gemma-4-31b, the gemini-* chat models, first-party IDs (gpt-4o, gpt-4.1, gpt-5-mini, gpt-5.5, gpt-5.6-*, gpt-6-*, gpt-6.1-sol, grok-4.3, DeepSeek-V4-*, gpt-realtime*, gpt-live-1, nova-sonic*, first-party STT/TTS), the other speech models (gnani-*, ink-*, sonic-3.6), the Deepgram direct line (deepgram-nova-3, deepgram-flux-general-en/multi, deepgram-nova-2*, deepgram-enhanced*, deepgram-base*, deepgram-whisper-*, deepgram-aura-2, deepgram-aura-1, Deepgram Voice Agent deepgram-voice-* ids, the deepgram-summarize/topics/sentiment/intents Audio Intelligence features, and the deepgram-text-summarize/topics/sentiment/intents Text Intelligence features), and paid image models (flux-2-pro, flux-1.1-pro, gpt-image-2.5-*, gpt-image-2, gpt-image-1.5, nano-banana-*, gemini-3.1-flash-lite-image) — require Starter, Pro, or Enterprise.
Pricing
All models are pay-per-use. Pricing is in US$ at US$1 = ₹96; you pay in credits, where 1 credit = ₹1 ≈ US$0.0104.
| Model | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
kimi-k2.5-fast | $0.8438 | $4.219 |
sarvam-105b | $0.3646 (₹35) | $0.3646 (₹35) |
sarvam-105b-conversations | $0.3646 (₹35) | $0.3646 (₹35) |
gpt-5.6-sol | $5.208 | $31.25 |
gpt-5.6-terra | $2.083 | $12.5 |
gpt-5.6-luna | $0.2083 | $1.25 |
gpt-6-sol | $2.083 | $10.42 |
gpt-6-luna | $0.1042 | $0.5208 |
gpt-6.1-sol | $2.083 | $10.42 |
nova-sonic-2 | $4.167 | $15.63 |
nova-sonic | $4.688 | $17.71 |
gpt-realtime | $4.167 | $16.67 |
gpt-realtime-mini | $0.625 | $2.5 |
gpt-realtime-2 | $4.167 | $25 |
gpt-realtime-1.5 | $4.167 | $16.67 |
gpt-realtime-2.1 | $4.167 | $25 |
gpt-realtime-2.1-mini | $0.625 | $2.5 |
gpt-live-1 | per-minute only | $0.05208/min (voice-agent only) |
deepgram-voice-* | per-minute Voice Agent tier | Standard $0.07813/min, Advanced $0.1698/min (voice-agent only) |
| STT Model | Price |
|---|---|
saaras:v3 | $0.3125 / hour (₹30/hr) |
saaras:v4 | $0.3125 / hour (₹30/hr) |
gnani-prisma-v2.5 | $0.2813 / hour |
ink-whisper | $0.1875 / hour |
ink-2 | $0.5625 / hour (voice sessions only) |
whisper-large-v3-turbo | $0.0625 / hour |
nova-3 | $0.5208 / hour |
deepgram-nova-3 | $0.3021 / hour |
deepgram-nova-3-medical | $0.3021 / hour |
deepgram-flux-general-en | $0.4063 / hour |
deepgram-flux-general-multi | $0.4896 / hour |
deepgram-nova-2 (+ domain variants) | $0.3646 / hour |
deepgram-nova / deepgram-whisper-* | $0.3646 / hour |
deepgram-enhanced (+ variants) | $1.031 / hour |
deepgram-base (+ variants) | $0.9063 / hour |
| TTS Model | Price |
|---|---|
bulbul:v3 | $0.3125 / 10K chars (₹30/10K) |
gnani-timbre-v2.0 | $0.2813 / 10K chars |
aura-2-en | $0.4167 / 10K chars |
aura-2-es | $0.4167 / 10K chars |
deepgram-aura-2 | $0.3125 / 10K chars |
deepgram-aura-1 | $0.1563 / 10K chars |
sonic-3.6 | $0.5208 / 10K chars |
melotts | $0.05208 / 10K chars |
| Intelligence feature (not a model ID) | Price |
|---|---|
deepgram-summarize / -topics / -sentiment / -intents (audio) | $0.0003125 / 1K input + $0.000625 / 1K output tokens |
deepgram-text-summarize / -text-topics / -text-sentiment / -text-intents | $0.0003125 / 1K input + $0.000625 / 1K output tokens |
These eight values go in the features field, not in model. They are not
catalog models — GET /api/v1/models/deepgram-summarize returns 404.
Full pricing for all models is available via the API: GET /api/v1/models
import requests
# List all LLM models
models = requests.get("https://api.callmissed.com/api/v1/models?category=llm").json()
for m in models["data"]:
print(f"{m['id']} — {m['name']} ({m['context_window']} tokens) {'FREE' if m['free'] else 'PAID'}")Fast LLMs
High-throughput Kimi K2.5 inference tier optimized for voice-agent latency.
| Model ID | Status | Context | Best For |
|---|---|---|---|
kimi-k2.5-fast | Under maintenance — fall back to kimi-k2.5 | 256K | Voice agents, fast inference, reasoning tasks |
While kimi-k2.5-fast is in maintenance (returns HTTP 503), use kimi-k2.5:
response = client.chat.completions.create(
model="kimi-k2.5",
messages=[{"role": "user", "content": "Hello"}]
)Indic Models
Speech to Text
| Model | Description | Languages |
|---|---|---|
saaras:v3 | Latest STT — best accuracy on Indian + code-mixed | 23 languages (22 Indic + English) |
saaras:v4 | Five output modes on one model — transcribe, translate, verbatim, transliterate, code-mix | 24 languages |
gnani-prisma-v2.5 | India-first telephony STT — code-switching, sub-4% WER on Indian English | 10 Indian languages |
For 99-language general-purpose transcription, see whisper-large-v3-turbo. For diarization + smart-format on calls, see nova-3. Both are free-tier and live under the audio model routes.
ink-whisper also covers Hindi, Urdu and Tamil as part of its 100-language set at $0.1875 / hr — cheaper than the Indic-specialist models, though without their code-mix output modes. ink-2 covers Hindi among its five languages (en, fr, hi, ja, es) but no other Indic language.
Text to Speech
| Model | Description | Voices |
|---|---|---|
bulbul:v3 | Natural TTS — 37 voices, 11 Indian languages | shubh (default) + 36 more |
gnani-timbre-v2.0 | India-first neural TTS — context-aware tone, low-latency. Served by Gnani's current Timbre generation (v2.5); the id keeps its original name | 73 voices (English + Hindi + Indic) |
sonic-3.6 | Cartesia Sonic 3.6 — most natural conversational speech, 44 languages with native-quality Hindi | Searchable public library + 16 featured aliases (skylar default) |
For low-latency English / Spanish voice agents, see aura-2-en / aura-2-es. For ultra-cheap en/fr notification audio, see melotts. All three are free-tier.
Chat Completion (LLM)
| Model | Params | Context | Best For |
|---|---|---|---|
sarvam-105b | 105B MoE | 128K tokens | Complex reasoning, agentic tasks, long documents |
sarvam-105b-conversations | 105B MoE | 32K tokens | Conversation and voice agents, tool calling |
glm-5.3 | Z.ai GLM 5.3 | 1M tokens | Long-context reasoning, tool calling (text only) |
gemma-4-31b | Google Gemma 4 31B | 128K tokens | Fast answers with thinking off, image input, tool calling, voice agents |
glm-5.3 always reasons. reasoning_effort takes "low", "high" or "max" (the default when omitted); "none" and "minimal" map to "low", "medium" to "high" and "xhigh" to "max". Reasoning tokens bill as output. gemma-4-31b answers without thinking by default. reasoning_effort: "low", "medium" or "high" switch thinking on; the level does not reliably change how much it thinks. "none" and "minimal" keep it off. Both are paid-plan models.
Both sarvam-105b models support hybrid thinking via reasoning_effort: "low" | "high" | "max", and "xhigh" maps to "max". "none" and "minimal" switch thinking fully off. "medium" is not one of their levels, so it is dropped and the model runs at its default. Other models with a full off switch: kimi-k2.5, kimi-k2.6, glm-4.7-flash, glm-5.2, gemma-4-26b-a4b-it, gemma-4-31b, DeepSeek-V4-Pro, DeepSeek-V4-Flash, or a GPT-5.5 / GPT-5.6 / GPT-6 model with reasoning_effort: "none". See the per-model matrix.
Audio Models
Free-tier on every plan. See the Pricing page for current rates.
Speech to Text
| Model | Languages | Best for | Price |
|---|---|---|---|
whisper-large-v3-turbo | 99 with auto-detect | Multilingual general-purpose; transcribe + translate | $0.0625 / hour |
nova-3 | 11 BCP-47 incl. multi auto-detect | Diarization, smart-format, streaming voice agents | $0.5208 / hour |
whisper | 99 with auto-detect | Whisper batch + translate | $0.4167 / hour |
gpt-4o-transcribe | Streaming | Higher-accuracy OpenAI transcription | $0.4167 / hour |
gpt-4o-mini-transcribe | Streaming | Low-cost OpenAI transcription | $0.25 / hour |
gpt-4o-transcribe-diarize | File transcription with speaker labels (not available for voice sessions) | Multi-speaker meetings / calls | $0.4167 / hour |
Deepgram (direct) — the full Deepgram speech-to-text line, billed per audio hour at the rates below:
| Model | Languages | Best for | Price |
|---|---|---|---|
deepgram-flux-general-en | English | Conversational voice agents — model-native turn detection, ultra-low latency | $0.4063 / hour |
deepgram-flux-general-multi | 10 (multilingual) | Multilingual voice agents with code-switching | $0.4896 / hour |
deepgram-nova-3 | 45+ incl. multi | Flagship general-purpose ASR, keyterm prompting, PII redaction | $0.3021 / hour |
deepgram-nova-3-medical | English | Clinical / medical terminology | $0.3021 / hour |
deepgram-nova-2 | 36 incl. multi | High-accuracy ASR + filler-word detection | $0.3646 / hour |
deepgram-nova-2-{meeting,phonecall,finance,conversationalai,voicemail,video,medical,drivethru,automotive,atc} | English | Domain-tuned Nova-2 variants | $0.3646 / hour |
deepgram-nova / -phonecall / -medical | en/es/hi | Legacy Nova-1 | $0.3646 / hour |
deepgram-enhanced (+ meeting/phonecall/finance) | 13 | Legacy, keyword boosting | $1.031 / hour |
deepgram-base (+ 6 variants) | 17 | Legacy, high-volume batch | $0.9063 / hour |
deepgram-whisper-{tiny,base,small,medium,large} | 99 | Whisper in five sizes, batch transcription | $0.3646 / hour |
Text to Speech
| Model | Languages | Voices | Price |
|---|---|---|---|
aura-2-en | English | 40 (luna default) | $0.4167 / 10K chars |
aura-2-es | Spanish | 10 (aquila default) | $0.4167 / 10K chars |
deepgram-aura-2 | en/es/de/fr/nl/it/ja | 90+ (thalia default) | $0.3125 / 10K chars |
deepgram-aura-1 | English | 12 (asteria default) | $0.1563 / 10K chars |
melotts | English + French | 1 per language | $0.05208 / 10K chars |
gpt-4o-mini-tts | Multilingual steerable | 6 OpenAI voices | $0.2083 / 10K chars |
Aura 2 returns linear16 PCM streamed at 24 kHz for low-latency playback. MeloTTS returns base64 MP3. Output formats may vary as models are updated.
Deepgram Flux TTS is a voice-agent-first model and is not available on this /v1/audio/speech endpoint. It is offered only through the managed Voice Agent (see the Voice Sessions API), selectable with tts_engine: "flux", where it is billed inside the per-minute voice rate.
Audio Intelligence (Deepgram)
Deepgram Audio Intelligence runs analysis over an uploaded audio file via POST /v1/audio/intelligence (English only, 150K input-token limit). Token-billed at $0.0003125/1K input + $0.000625/1K output.
| Feature | Model ID | Returns |
|---|---|---|
| Summarization | deepgram-summarize | A concise summary of the audio |
| Topic Detection | deepgram-topics | Per-segment topics with confidence |
| Sentiment Analysis | deepgram-sentiment | Per-segment + average sentiments |
| Intent Recognition | deepgram-intents | Per-segment intents with confidence |
Request multiple features in one call with a comma-separated features form field (e.g. features=deepgram-summarize,deepgram-sentiment).
Text Intelligence (Deepgram)
Deepgram Text Intelligence runs the same four analyses over text input (a string or a hosted text URL) via POST /v1/text/intelligence (English only, 150K input-token limit). Token-billed at $0.0003125/1K input + $0.000625/1K output. Requires the llm key permission.
| Feature | Model ID | Returns |
|---|---|---|
| Summarization | deepgram-text-summarize | A concise summary of the text |
| Topic Detection | deepgram-text-topics | Per-segment topics with confidence |
| Sentiment Analysis | deepgram-text-sentiment | Per-segment + average sentiments |
| Intent Recognition | deepgram-text-intents | Per-segment intents with confidence |
Send a JSON body with features (array or comma-separated string) and exactly one of text or url:
{
"features": ["deepgram-text-summarize", "deepgram-text-sentiment"],
"text": "Your text to analyze here."
}Direct-Routed LLMs
Low-latency models routed directly through CallMissed — sub-2s end-to-end on small prompts and free-tier eligible per the reasoning_effort matrix.
| Model ID | Creator | Context |
|---|---|---|
kimi-k2.5 | Moonshot AI | 256K |
kimi-k2.6 | Moonshot AI | 262K |
kimi-k2.7-code | Moonshot AI | 262K |
gpt-oss-120b | OpenAI (open-weights) | 128K |
gemma-4-26b-a4b-it | 256K | |
glm-4.7-flash | Zhipu | 128K |
glm-5.2 | Z.ai | 262K |
nemotron-3-super | NVIDIA | 256K |
mistral-small-3.1 | Mistral | 128K |
Models on Demand
GET /api/v1/models lists everything in the catalogue today: 138 model IDs.
Models marked status: "maintenance" return 503 until they are back.
Beyond that we deploy 300+ further models on demand on CallMissed
infrastructure. Send the model you need and your expected throughput to
sales@callmissed.com. Once deployed it appears in your GET /api/v1/models
response with a plain CallMissed ID and published per-token pricing, on the
same /v1/chat/completions endpoint as every other model. Same key, same
credit balance, no new SDK.
Enterprise accounts get dedicated capacity. Starter and Pro get shared capacity where the model allows it.
First-Party Models
Credit-covered first-party models. Use the bare model ID in API requests — e.g. gpt-4o.
| Model ID | Type | Notes |
|---|---|---|
gpt-4o | LLM | Multimodal text + vision, 128K context |
gpt-4.1 | LLM | Long-context (300K) multimodal |
gpt-5-mini | LLM | Fast reasoning, 400K context |
gpt-5.5 | LLM | GPT-5.5 reasoning flagship, 1.05M context, vision + tools |
gpt-5.6-sol | LLM | GPT-5.6 flagship, 1.05M context, vision + tools |
gpt-5.6-terra | LLM | GPT-5.6 balanced intelligence/cost, 1.05M context |
gpt-5.6-luna | LLM | GPT-5.6 fast + affordable, 1.05M context |
gpt-6-sol | LLM | GPT-6 frontier reasoning, 1.05M context, vision + tools |
gpt-6-luna | LLM | GPT-6 efficient, high-volume, 1.05M context |
gpt-6.1-sol | LLM | GPT-6.1 near-frontier for complex coding and professional work, 1.05M context, vision + tools |
grok-4.3 | LLM | xAI Grok, 200K context |
DeepSeek-V4-Pro | LLM | Flagship DeepSeek reasoning, 1M context, tools (text-only) |
DeepSeek-V4-Flash | LLM | Fast DeepSeek reasoning, 1M context, tools |
nova-sonic-2 | Realtime voice | Amazon Nova 2 speech-to-speech voice model — 16 voices, Hindi + en-IN, live |
nova-sonic | Realtime voice | First-generation Amazon speech-to-speech voice model — retired upstream 2026-09-14, use nova-sonic-2 |
gpt-realtime | Realtime voice | OpenAI flagship speech-to-speech (10 concurrent), live |
gpt-realtime-mini | Realtime voice | Lowest-cost realtime, ~3× cheaper than gpt-realtime, live |
gpt-realtime-2 | Realtime voice | Newest realtime with stronger tool calling, live |
gpt-realtime-1.5 | Realtime voice | Pinned 1.5 snapshot of gpt-realtime, live |
gpt-realtime-2.1 | Realtime voice | Latest realtime — better recognition, silence/interrupt handling, live |
gpt-realtime-2.1-mini | Realtime voice | Distilled low-cost 2.1 realtime, live |
gpt-live-1 | Realtime voice | GPT Live speech-to-speech — audio + text only, 14 voices, $0.05208/min, live |
whisper | STT | OpenAI Whisper — 99 langs |
gpt-4o-transcribe | STT | Streaming transcription |
gpt-4o-mini-transcribe | STT | Low-cost streaming STT |
gpt-4o-transcribe-diarize | STT | Speaker diarization |
gpt-4o-mini-tts | TTS | Steerable OpenAI TTS, 13 voices |
See Credits & Rate Limits for per-model USD pricing.
Full Model Catalog
A curated, representative slice of the 138 models (67 LLM · 45 STT · 9 TTS · 15 image · 2 embedding) served by GET /api/v1/models as of the latest deploy — the per-domain variants of the direct Deepgram STT line and the deepgram-voice-* managed LLM ids are covered in their own sections above rather than repeated below. For live pricing and capability flags (supports_vision, supports_tools, free), query the API — it always reflects the current catalog.
LLM (43 models)
| Model ID | Description | Context | Free | Pricing |
|---|---|---|---|---|
sarvam-105b | 105B MoE. Complex reasoning, agentic tasks, long documents. | 131K | Yes | $0.3646 in / $0.3646 out per 1M |
sarvam-105b-conversations | 105B MoE tuned for conversation and voice. Tool calling. | 32K | Yes | $0.3646 in / $0.3646 out per 1M |
gpt-4o | Multimodal text + vision. | 128K | No | $2.604 in / $10.42 out per 1M |
gemini-3.8-flash | Fast multimodal flagship. Thinking low/medium/high. | 1M | No | $1.563 in / $7.813 out per 1M |
gemini-3.7-flash | Fast multimodal. Thinking low/medium/high. | 1M | No | $1.563 in / $7.813 out per 1M |
gemini-3.6-flash | Fast multimodal. Thinking minimal→high. | 1M | No | $1.563 in / $7.813 out per 1M |
gemini-3.5-flash | Balanced multimodal workhorse. | 1M | No | $1.563 in / $9.375 out per 1M |
gemini-3.5-flash-lite | Cheapest 1M-context Gemini. | 1M | No | $0.3125 in / $2.604 out per 1M |
gemini-3.1-pro-preview | Reasoning-heavy Gemini tier. | 1M | No | $2.083 in / $12.5 out per 1M |
gemini-3.1-flash-lite | Low-cost multimodal, tool use. | 1M | No | $0.2604 in / $1.563 out per 1M |
gpt-4.1 | Long-context multimodal. Strong instruction following. | 300K | No | $2.083 in / $8.333 out per 1M |
gpt-5-mini | Fast, affordable reasoning. | 400K | No | $0.2604 in / $2.083 out per 1M |
gpt-5.5 | Reasoning flagship. Vision, tools, prompt caching. | 1.05M | No | $5.208 in / $31.25 out per 1M |
gpt-5.6-sol | Frontier model for complex professional work. Vision, reasoning, tools. | 1.05M | No | $5.208 in / $31.25 out per 1M |
gpt-5.6-terra | Balances intelligence and cost. Vision, reasoning, tools. | 1.05M | No | $2.083 in / $12.5 out per 1M |
gpt-5.6-luna | Cost-sensitive, high-volume workloads. Vision, reasoning, tools. | 1.05M | No | $0.2083 in / $1.25 out per 1M |
gpt-6-sol | Frontier reasoning for enterprise agents, coding and complex knowledge work. Vision, reasoning, tools. | 1.05M | No | $2.083 in / $10.42 out per 1M |
gpt-6-luna | Efficient GPT-6 for high-volume, cost-sensitive workloads. Vision, reasoning, tools. | 1.05M | No | $0.1042 in / $0.5208 out per 1M |
gpt-6.1-sol | Near-Astra performance for complex coding, computer use and professional work at a lower cost. Vision, reasoning, tools. | 1.05M | No | $2.083 in / $10.42 out per 1M |
grok-4.3 | xAI Grok 4.3. Reasoning + vision (images of at least 512 pixels in total; smaller returns 400). Reasoning tokens bill as output. | 200K | No | $3.646 in / $15.63 out per 1M |
DeepSeek-V4-Pro | Flagship DeepSeek reasoning. Tools. Text-only. | 1M | No | $1.375 in / $4.125 out per 1M |
DeepSeek-V4-Flash | Fast, affordable DeepSeek reasoning. Tools. | 1M | No | $0.4583 in / $1.375 out per 1M |
kimi-k2.5 | Strong on coding and math. Vision. | 256K | Yes | $0.8438 in / $4.219 out per 1M |
kimi-k2.5-fast (maintenance) | Kimi K2.5 at ~414 tok/s for voice-agent latency. | 256K | No | $0.8438 in / $4.219 out per 1M |
kimi-k2.6 | Improved reasoning and coding over K2.5. Vision. | 262K | Yes | $1.333 in / $5.625 out per 1M |
kimi-k2.7-code | 1T-param agentic coding. Vision + tools. | 262K | Yes | $1.333 in / $5.625 out per 1M |
glm-4.7-flash | Fast, cost-efficient bilingual model. Strong tool use. | 131K | Yes | $0.5208 in / $2.083 out per 1M |
glm-5.2 | Flagship agentic coding. Tools + reasoning. | 262K | Yes | $1.969 in / $6.188 out per 1M |
glm-5.3 | Z.ai GLM 5.3. Tools + reasoning (low / high / max). Text-only. | 1M | No | $2.125 in / $6.667 out per 1M |
gpt-oss-120b | Open-weight 120B MoE. Reasoning-grade at lower cost. | 128K | Yes | $1.042 in / $4.167 out per 1M |
nemotron-3-super | 120B MoE tuned for long-context reasoning. | 256K | Yes | $1.563 in / $6.25 out per 1M |
gemma-4-26b-a4b-it | 26B MoE (4B active). Efficient instruct model. Vision. | 256K | Yes | $0.4167 in / $1.667 out per 1M |
gemma-4-31b | Gemma 4 31B instruct. Fast with thinking off, vision, tools. | 128K | No | $0.6146 in / $1.542 out per 1M |
mistral-small-3.1 | 24B instruct. Strong tool use, fast. Vision. | 128K | Yes | $0.4896 in / $0.7917 out per 1M |
nova-sonic-2 | Amazon Nova 2 Sonic. Native speech-to-speech voice model — STT, reasoning, and TTS in one; 16 voices across 8 languages including Hindi + en-IN. | 1M | No | $4.167 in / $15.63 out per 1M • $0.06667/min |
nova-sonic | Amazon Nova Sonic 1.0. Native speech-to-speech voice model with 11 voices across English, Spanish, French, Italian, and German. | 32K | No | $4.688 in / $17.71 out per 1M • $0.07396/min |
gpt-realtime | OpenAI flagship realtime speech-to-speech model — STT + reasoning + function calling + TTS in one. 10 concurrent. | 32K | No | $4.167 in / $16.67 out per 1M • $0.3906/min |
gpt-realtime-mini | Lowest-cost realtime — same single-model shape as gpt-realtime, ~3× cheaper. 20 concurrent. | 32K | No | $0.625 in / $2.5 out per 1M • $0.1229/min |
gpt-realtime-2 | Newest realtime with stronger tool calling. 32K input, 4K output. | 32K | No | $4.167 in / $25 out per 1M • $0.3906/min |
gpt-realtime-1.5 | Pinned 1.5 snapshot of gpt-realtime. Use when you want version stability. | 32K | No | $4.167 in / $16.67 out per 1M • $0.3906/min |
gpt-realtime-2.1 | Latest realtime speech-to-speech — better alphanumeric recognition, silence/noise + interruption handling. Voice-agent only. | 32K | No | $4.167 in / $25 out per 1M • $0.3906/min |
gpt-realtime-2.1-mini | Distilled, lower-cost realtime for faster voice interactions. Voice-agent only. | 32K | No | $0.625 in / $2.5 out per 1M • $0.1219/min |
gpt-live-1 | GPT Live speech-to-speech — audio + text in and out, function calling, 14 voices (default marin). No image or video input. Voice-agent only. | — | No | $0.05208/min (billed per second) |
Speech to Text (11 models)
| Model ID | Description | Context | Free | Pricing |
|---|---|---|---|---|
saaras:v3 | 23 languages (22 Indic + English). Best on code-mixed speech. | — | Yes | $0.3125 / hr |
saaras:v4 | 24 languages. Five output modes: transcribe, translate, verbatim, transliterate, code-mix. | — | Yes | $0.3125 / hr |
gnani-prisma-v2.5 | India-first telephony STT. 10 Indian languages, code-switching. | — | No | $0.2813 / hr |
ink-whisper | Cartesia Ink Whisper — 100 languages including Hindi, Urdu and Tamil. Dynamic chunking reduces hallucination across pauses and silence. File transcription + streaming. | — | No | $0.1875 / hr |
ink-2 | Cartesia Ink 2 — top-ranked for voice agents (8% WER on AppTek's 14-accent call-centre benchmark, vs 10% Deepgram Flux and 12% ElevenLabs). Self-detects turns. Languages: en, fr, hi, ja, es. Voice sessions only — not available for file transcription. | — | No | $0.5625 / hr |
whisper-large-v3-turbo | 99 languages with auto-detect. Transcribe + translate. | — | Yes | $0.0625 / hr |
nova-3 | Diarization, punctuation, smart-format. Streaming-capable. | — | Yes | $0.5208 / hr |
whisper | 99 languages. Transcription + translation to English. | — | No | $0.4167 / hr |
gpt-4o-transcribe | Higher accuracy than Whisper. Streaming. | — | No | $0.4167 / hr |
gpt-4o-mini-transcribe | Cheaper, faster streaming transcription. | — | No | $0.25 / hr |
gpt-4o-transcribe-diarize | File transcription with speaker labels. Not available for voice sessions. | — | No | $0.4167 / hr |
Text to Speech (7 models)
| Model ID | Description | Voices | Free | Pricing |
|---|---|---|---|---|
bulbul:v3 | Indic TTS across 11 Indian languages. | 37 | Yes | $0.3125 / 10K chars |
gnani-timbre-v2.0 | India-first neural TTS, English + Hindi + Indic. Context-aware tone. Served by Timbre v2.5. | 73 | No | $0.2813 / 10K chars |
sonic-3.6 | Cartesia Sonic 3.6 — most natural conversational TTS. 44 languages, native-quality Hindi + Hinglish, sub-90ms first audio. | Searchable library | No | $0.5208 / 10K chars |
aura-2-en | Conversational English TTS, low-latency streaming. | 40 | Yes | $0.4167 / 10K chars |
aura-2-es | Spanish TTS, low-latency streaming. | 10 | Yes | $0.4167 / 10K chars |
melotts | Lightweight English + French TTS. Cheapest available. | 1 per language | Yes | $0.05208 / 10K chars |
gpt-4o-mini-tts | Steerable — takes an instructions field to direct tone. | 6 | No | $0.2083 / 10K chars |
Image Generation (15 models)
| Model ID | Description | Free | Pricing |
|---|---|---|---|
flux-2-klein-9b | Flux 2 Klein. 1024×1024 default. | Yes | $0.1042 / image |
flux-2-dev | Flux 2 Dev. Higher fidelity, 50-step inference. | Yes | $0.125 / image |
flux-2-pro | Flux 2 Pro. Flagship BFL fidelity. | No | $0.1042 / image |
flux-1.1-pro | Flux 1.1 Pro. Fast, production-grade. | No | $0.05208 / image |
gpt-image-2.5-sunburst | Most capable generation + editing. Inpainting, quality tiers to max. | No | $0.2604 / image |
gemini-3.1-flash-lite-image | Low-cost text-to-image with reference edits. 1K resolution. | No | $0.035 / image |
gpt-image-2.5-flare | Fast, high-quality everyday generation. On-image text + edits. | No | $0.2604 / image |
gpt-image-2 | Accurate on-image text rendering. | No | $0.2604 / image |
gpt-image-1.5 | Precise image editing. Strong logo/face preservation. | No | $0.2604 / image |
lucid-origin | Vibrant, cinematic compositions. | Yes | $0.08333 / image |
phoenix-1.0 | Strong prompt adherence, photorealistic portraits. | Yes | $0.1042 / image |
sdxl-lightning | 4-step inference. Fastest for iterative prompting. | Yes | $0.04167 / image |
dreamshaper-8-lcm | Stylised illustrations, fast generation. | Yes | $0.04167 / image |
nano-banana-2 | Fast multimodal image generation. | No | $0.06979 / image |
nano-banana-pro | Flagship typography and fidelity. | No | $0.1396 / image |
Embeddings (2 models)
| Model ID | Description | Dimensions | Free | Pricing |
|---|---|---|---|---|
text-embedding-3-small | Fast, low-cost embeddings. Best price/performance for large corpora. | 1536 | Yes | $0.02083 / 1M input tokens |
text-embedding-3-large | Highest-accuracy embeddings. | 3072 | Yes | $0.1354 / 1M input tokens |
Both accept 8,192-token inputs and support shortening the vector with dimensions. See Embeddings.
Tip: Filter programmatically —
GET /api/v1/models?category=llm,?category=stt,?category=tts,?category=image,?category=embedding, or?free=truefor free-plan models only.
Model Selection
Pass the model ID in your request:
# Indic LLM
response = client.chat.completions.create(
model="sarvam-105b",
messages=[{"role": "user", "content": "Hello in Hindi"}]
)
# Indic LLM with thinking mode
response = client.chat.completions.create(
model="sarvam-105b",
messages=[{"role": "user", "content": "Solve this step by step"}],
extra_body={"reasoning_effort": "high"}
)
# First-party flagship model
response = client.chat.completions.create(
model="gpt-5.6-luna",
messages=[{"role": "user", "content": "Hello"}]
)The API automatically routes to the correct backend based on the model ID:
- Bare names (
kimi-k2.5,gpt-4o,DeepSeek-V4-Pro,mistral-small-3.1, …) → direct-routed or first-party sarvam-*prefix → Indic LLMssaaras:*/bulbul:*/deepgram-*/ image IDs → the matching speech or image backend
Every ID is a plain CallMissed ID with no vendor prefix — including models we deploy on demand.