Changelog
Latest updates, new features, and improvements to the CallMissed API.
October 2026
Usage API — costs in US$, prompt-cache totals
cost_usdon/v1/usage/summary,/v1/usage/logsand/v1/usage/logs.csv(and theget_usage_summary/list_usage_logsMCP tools) is now the US$ you were billed at the published rate (1 credit = ₹1, US$1 = ₹96). It was previously reported in internal billing units (0.01per credit), about 4% below the US$ amount; credits = US$ × 96. Past rows are converted the same way, so a window spanning the change stays consistent./v1/usage/summaryreports prompt-cache usage:totals.total_cache_read_tokens,totals.total_cache_creation_tokens, andcache_read_tokens/cache_creation_tokensper service.input_tokenskeeps its meaning (the uncached part of the prompt).
Per-model parameters follow each model's documentation
- Output limits —
GET /api/v1/modelsnow listsmax_output_tokenswhere a model documents one, andmax_tokensabove it returns400 max_tokens_too_largeon/v1/chat/completions,/v1/responsesand/v1/messages. Thegemini-*models now allow up to 65,536 output tokens (previously capped at 8,192), and an omittedmax_tokensleaves the model's own default. reasoning_effort—xhighandmaxnow reach every model instead of being cut tohighfirst; each model maps them to its own top level.DeepSeek-V4-Pro/-Flashandglm-5.2now honour their levels,nemotron-3-superswitches to its low-effort mode forlow/none,sarvam-105bswitches thinking fully off fornone, andgpt-5-miniusesminimal. See API speed.- Sarvam — when you omit
max_tokens, we send 4,096 (16,384 forglm-5.3) so thinking no longer cuts answers short at 2,048. Aglm-5.3orgemma-4-31brequest over 10 MB returns413. - Models —
gpt-4.1serves a 300,000-token context window. The realtimegpt-realtime-2*models take 32,000 input and 4,096 output tokens.nova-sonicwas retired upstream on 2026-09-14 and returns503; usenova-sonic-2. - Images and speech — image sizes outside a model's range, and
negative_promptonlucid-origin, return400.gpt-4o-mini-ttshas 13 voices and rejects an unknown voice with400.gpt-4o-transcribe-diarizereturns speakersegmentswithresponse_formatverbose_jsonordiarized_json; it is no longer accepted as a voice-sessionstt_model(422), because live transcription cannot label speakers. grok-4.3 reasoning tokens now count as output tokens, and its images must have at least 512 pixels.
Email domains — neutral 502 reason
- The
502returned by the email domain routes when provisioning or verification is temporarily unavailable now carriesreason: "provider_unavailable". Deployments still being upgraded may return the olderacs_unavailablefor the same condition; status, body shape and retry advice are identical, so match both strings. See Email limits.
Voice session webhooks
voice_session.endedandvoice_session.failednow fire exactly once per session, on every way a session can end (the agent hanging up,DELETE, the session timing out), and only after the end has been saved. If two end signals arrive together, the first one decidesend_reasonandduration_seconds.- A session that ends because the agent hit an error (
end_reason: "agent_error") now has statusfailedand firesvoice_session.failedon every path; it previously firedvoice_session.endedfor some sessions. See Webhooks.
Speech-to-text error codes
/v1/audio/transcriptionsand/v1/audio/translationsnow return400 invalid_requestfor a problem with the request (unsupported audio, a streaming-only or unknown model),429 rate_limit_exceededwhen rate limited, and502 service_unavailablewhen transcription is temporarily unavailable, the same code/v1/audio/speechuses. See Speech-to-Text.
Batch API, zero data retention, per-end-user budgets and prompt caching
- Batch API — OpenAI-compatible
POST /v1/filesandPOST /v1/batches: upload a JSONL file of requests, run it asynchronously, and download the results. Gated by the key'sllmpermission. See Batch API. - Zero data retention — send
"provider": {"zdr": true}to route a request only to models that keep no data, or enforce it for a whole key or account. See Gateway controls. - Per-end-user budgets — monthly credit caps per end user of your app, managed under
/api/v1/gateway/end-user-budgets(scopesend_user_budgets:read/end_user_budgets:write). See Gateway controls. - Prompt caching —
prompt_cache_keyon/v1/chat/completionsand/v1/responsespasses through as a cache-routing hint to the models that support it (and is dropped elsewhere);cache_controlblocks on/v1/messagesare accepted so Anthropic SDK code runs unchanged, but explicit breakpoints are not applied yet — caching works from the prompt prefix.usage.prompt_tokens_details.cached_tokensis now always present, and the Usage API reports cache reads and writes.
Payments, evals, contacts and telephony
- Customer payment links — agents can send payment links on calls and WhatsApp that are paid into your own connected Razorpay account. Track them under
/api/v1/payment-requestsand subscribe to thepayment_request.*webhook events. - Evals — an
llm_judgeassertion, a pass-rate gate for CI, and creating an eval case from a real call. See Voice evals. - WhatsApp identities on contacts — contacts carry
whatsapp_bsuid,whatsapp_parent_bsuidandwhatsapp_username, for WhatsApp users who hide their phone number. A BSUID is unique per account like a phone number or email. See Contacts. - TRAI readiness for outbound campaigns — a readiness check for AI outbound calling under the TCCCPR amendment. See Voice campaigns.
Newly documented
- Contacts (
/api/v1/contacts, plus each contact's memory), the handoff queue (/api/v1/handoffs), integrations and sheet automations (/api/v1/integrations,/api/v1/sheet-automations), payment requests, andPOST /api/v1/crm/csv/inspectwere already callable with an API key and now have reference pages. - Webhooks — the delivery body, the
X-CallMissed-Event/X-CallMissed-Deliveryheaders, the retry schedule, and thecall.amd_detectedevent. - Idempotency and Errors — rewritten to match exactly what the API does.
September 2026
Account MCP server — run your account from an AI assistant
- Account MCP server — connect Claude, ChatGPT or any MCP client to
https://api.callmissed.com/api/v1/mcp, sign in with your CallMissed account or an API key, and act on your account through 331 tools: voice agents, calls and campaigns, the inbox, CRM, support desk, WhatsApp, email and more. Every tool is gated by the key's scopes, spending tools are marked, and destructive ones ask first. See Account MCP. - Multilingual voice agents — voice tiers and custom agents accept an auto-detect language, so one agent can answer in the caller's language. See Voice tiers.
New model — GPT-6.1 Sol
gpt-6.1-sol— OpenAI's GPT-6.1 Sol: near-Astra performance for complex coding, computer use and professional work at a lower cost. 1.05M context, 128K output, text + image input, reasoning and tool calling. $2.083 in / $10.42 out per 1M tokens ($0.1042 cached input).- Long prompts — requests with more than 272K input tokens are billed at 2× input and cached-input rates and 1.5× output for the whole request ($4.167 in / $0.2083 cached / $15.63 out per 1M).
reasoning_effort— acceptslow,medium(default),highandxhigh. The model always reasons, sononeandminimalare sent aslow, andmaxis sent asxhigh. Only the default temperature is supported, andmax_tokensmust be at least 3. See API speed.- Paid plans only. See Models.
New models — GPT-6 Sol and GPT-6 Luna
gpt-6-sol— OpenAI's GPT-6 frontier reasoning model for enterprise agents, coding and complex knowledge work; succeedsgpt-5.6-sol. 1.05M context, 128K output, text + image input, reasoning and tool calling. $2.083 in / $10.42 out per 1M tokens ($0.2083 cached input).gpt-6-luna— the efficient GPT-6 model for high-volume, cost-sensitive workloads; succeedsgpt-5.6-luna. Same 1.05M context, vision, reasoning and tools. $0.1042 in / $0.5208 out per 1M tokens ($0.01042 cached input).- Long prompts — requests with more than 272K input tokens are billed at 2× input and cached-input rates and 1.5× output for the whole request.
reasoning_effort— both acceptnone,low,medium,highandxhigh(minimalis sent aslow). When a request includestools, reasoning is set tononeautomatically.max_tokensmust be at least 3. See API speed.- Paid plans only. See Models.
August 2026
Managed Voice Agent — speech-to-speech over one WebSocket
- Managed Voice Agent — a full speech-to-speech pipeline behind a single WebSocket. Stream microphone audio in, get synthesized speech and conversation events back; speech recognition, the language model, text-to-speech, turn-taking and interruption handling are all run and tuned for you. No WebRTC and no client SDK. Two protocols on the same host and the same engine:
wss://api.callmissed.com/v2/voice/agent(CallMissed-native) andwss://api.callmissed.com/v1/agent/converse(Deepgram Voice Agent compatible — an existing Deepgram integration can repoint its URL and work unchanged). Supports in-call tool calling, live model/prompt/voice updates, and sessions up to 2 hours. See Managed Voice Agent. - Voice model catalogue —
GET /api/v1/voice/modelslists every selectable speech-to-text, language and text-to-speech model with a measured latency verdict (eligible,too_slow,unsupported,unmeasured) plus its p50, sample count and budget. Any eligible combination is valid. Models too slow to hold a conversation are not offered — they are listed with the reason, rather than silently disappearing or being offered with a caveat. Verdicts come from real production traffic, not vendor claims, and a model with too few samples readsunmeasuredrather than being assumed fast.
Embeddings, usage API, CRM, support desk and voice-agent operations
- Embeddings —
POST /v1/embeddings, OpenAI-compatible.text-embedding-3-small(1536 dims, $0.02083 / 1M input tokens) andtext-embedding-3-large(3072 dims, $0.1354 / 1M). Batches of up to 128 inputs, optionaldimensionsshortening andbase64output. Both are free-plan callable, taking the free tier to 27 models across five categories. Gated by the key'sllmpermission. See Embeddings. - Usage API —
GET /v1/usage/summary,/logsand/logs.csvreturn your own metering rows for the last 90 days, filterable by service, model, key,session_idandtrace_id. Scopeusage:read. See Usage API. - Gateway tooling — server-side prompt management with versions, labels, presets and free rendering; response cache stats and purge; and bring your own provider key with liveness verification and a write-only secret.
- CRM — companies, notes and tasks, deals and pipelines, custom fields and saved views, search, bulk and CSV, and lead scoring with a unified timeline.
- Support desk — tickets with server-managed lifecycle stamps, SLA policies with business hours and live breach reporting, macros, tags and routing rules with a dry-run evaluator, and CSAT/NPS surveys and a hosted page where customers answer.
- Voice-agent operations — eval suites (up to 50 cases per run, credit-charged), A/B experiments with deterministic assignment, and agent squads with handoff simulation and credit-charged agent drafting.
- WhatsApp — Flows (create, publish, read submissions) and catalog orders.
New models — Cartesia Ink STT
ink-whisper— Cartesia's fastest and most affordable STT at $0.1875 / hour, across 100 languages including Hindi, Urdu and Tamil. Better accuracy than baseline Whisper, and dynamic chunking that cuts hallucination during pauses and silence. Works for both file transcription and voice sessions. See Speech to Text.ink-2— Cartesia's top-ranked STT for voice agents at $0.5625 / hour: 8% WER on AppTek's 14-accent call-centre benchmark, against 10% for Deepgram Flux and 12% for ElevenLabs. Self-detects turns, so no separate turn detector is needed. Two limits: it is English only, and it is voice-session only — the file transcription endpoint returns400and points you toink-whisper. See Speech to Text.
New models — conversational Indic LLM, Saaras V4 STT, Flux TTS
sarvam-105b-conversations— 105B MoE tuned for conversation and voice. 128K context, tool calling, streaming, hybrid thinking. Free-tier, same $0.3646 in / $0.3646 out per 1M assarvam-105b. See Indic Models.saaras:v4— Sarvam STT with five output modes (transcribe, translate, verbatim, transliterate, code-mix) across 24 languages. Free-tier at $0.3125 / hour. See Speech to Text.- Deepgram Flux TTS — streaming-first TTS built for voice agents: turn-based synthesis with prosody carried across turns. 36 English voices including
priya(Indian-accented English, the default). English only, no expressive controls. Available only through the managed Voice Agent (tts_engine: "flux"), billed inside the per-minute voice rate. See Voices. - Free tier — now 27 models (11 LLM, 4 STT, 4 TTS, 6 image, 2 embedding).
Model catalog update — retired models
- Retired LLM IDs — the following model IDs are no longer served:
openai/gpt-5.4-pro,openai/gpt-5.4,openai/gpt-5.4-mini,openai/gpt-5.4-nano,anthropic/claude-opus-4.6,anthropic/claude-sonnet-4.6,anthropic/claude-haiku-4.5,x-ai/grok-4.20,qwen/qwen3.5-plus,qwen/qwen3.5-flash,mistralai/mistral-small-2603, and theautoauto-router. - Migration — use the first-party flagships (
gpt-5.6-sol,gpt-5.6-terra,gpt-5.6-luna,gpt-5.5,grok-4.3) or the direct-routed free tier (kimi-k2.6,kimi-k2.7-code,glm-5.2,gpt-oss-120b,mistral-small-3.1). See Models. - Free tier — now 24 models (11 LLM). The
autofree auto-router is retired; pick a free model explicitly. - Endpoints unchanged —
POST /v1/chat/completionsand the Anthropic-compatiblePOST /v1/messagesboth continue to work and accept every current catalog ID.
June 2026
v1.6.0 — WebRTC Voice, Image Generation & Web Search
- WebRTC voice sessions —
POST /v1/voice/sessionsreturns a connection token + URL; CallMissed handles the STT→LLM→TTS pipeline. List, fetch, fetch transcript (json | txt | srt), and end sessions under/v1/voice/sessions. The legacy/ws/voice-agentWebSocket still works for backward compatibility. See Voice Session API. - Image Generation API —
POST /v1/images/generations(OpenAI-compatible). Free models includeflux-2-klein-9b,flux-2-dev,lucid-origin,phoenix-1.0,sdxl-lightning,dreamshaper-8-lcm; paid models includeflux-2-pro,flux-1.1-pro,nano-banana-2, andnano-banana-pro. See Image Generation. - Web Search API —
POST /v1/searchdefaults to Serper web search (Exa, Firecrawl also available); flat 1 credit per query. See Web Search. - Knowledge RAG — vector knowledge sources at
/api/v1/knowledge/sources(ingest text, URL, or PDF; chunked + embedded) with semantic search atPOST /api/v1/knowledge/search.
May 2026
v1.5.0 — First-Party Models, Account Security & WhatsApp Platform
- First-party models — deployments callable by bare ID:
gpt-4o,gpt-4.1,gpt-5-mini,grok-4.3,DeepSeek-V4-Pro,DeepSeek-V4-Flash, plus first-party STT (whisper,gpt-4o-transcribe,gpt-4o-mini-transcribe,gpt-4o-transcribe-diarize) and TTS (gpt-4o-mini-tts). - More STT/TTS —
whisper-large-v3-turbo(99 langs),nova-3(diarization),aura-2-en/aura-2-es, andmelotts— all free-tier. - Kimi K2.6 —
kimi-k2.6added to the direct-routed free tier alongsidekimi-k2.5. - TOTP 2FA & passkeys — two-factor auth (authenticator apps + backup codes) and passkeys for dashboard sign-in. Active sessions can be reviewed and revoked from the dashboard.
- WhatsApp platform — Embedded Signup onboarding, message templates, broadcast campaigns, and delivery analytics under
/api/v1/whatsapp/*. See WhatsApp API. - Billing surfaces — coupon redemption, downloadable PDF invoices, and a credit ledger broken down by transaction type, all in the dashboard.
- Audit log — a sensitive-action audit feed in the dashboard.
April 2026
v1.4.0 — Anthropic API Compatibility & Audio Translation
- Anthropic Messages API — New
POST /v1/messagesendpoint. Use the Anthropic SDK with CallMissed by changing only thebase_url. Full streaming support with Anthropic SSE lifecycle (message_start,content_block_delta,message_stop). - Dual auth headers — Anthropic endpoint accepts both
x-api-keyandAuthorization: Bearerheaders - Model aliasing — A bare model name on the Anthropic endpoint resolves against the CallMissed catalog
- Audio Translation — New
POST /v1/audio/translationsendpoint. Translate audio in 24 languages to English text. OpenAI SDK compatible (client.audio.translations.create()) - Token counting —
POST /v1/messages/count_tokensfor input token estimation - Anthropic rate limit headers —
anthropic-ratelimit-requests-limit,anthropic-ratelimit-requests-remaining, etc.
v1.3.0 — Voice Agent & Ultra-Low-Latency Pipeline
- Voice Agent WebSocket — Real-time STT→LLM→TTS pipeline over
/ws/voice-agent. PCM audio in, streaming MP3 out. LLM and TTS run concurrently for minimum latency. - PCM AudioWorklet capture — Raw PCM s16le at 16kHz, no container overhead
- Streaming MP3 playback — MediaSource API appends and plays chunks as they arrive
- Profile management — Save and update user profile from the dashboard
v1.2.0 — Security, Google OAuth & Plan Enforcement
- Sign in with Google — Google sign-in for the dashboard. Auto-creates the organisation and user, and links to an existing account by email.
- OTP Authentication — Email-based OTP for passwordless login and password reset
- Plan limit enforcement — Server-side usage caps per plan tier (free/starter/pro/enterprise). API returns
429 quota_exceededwhen limits reached. Usage headers (X-RateLimit-*,X-Usage-Warning) on every response. - Per-API-key rate limiting — a requests-per-minute limit on every key (today set by plan; see Rate limits)
- Model catalog update — OpenAI gpt-5.4 family, Anthropic Claude 4.6, Google Gemini 3.1, xAI Grok 4.20, Qwen 3.5, Mistral Small
- Knowledge Base file upload — Upload PDF, DOCX, TXT files (max 20 MB) with auto text extraction
- Bot deployment verification — Verify WhatsApp/Twilio channel connectivity from the dashboard
- Settings verification — Verify WhatsApp, Twilio, and Indic LLM API connectivity
v1.1.0 — Platform Playground
- Playground rebuild — LLM (streaming + non-streaming), STT (file upload + mic), TTS (37 voices across 11 Indian languages), Voice Agent demo
v1.0.0 — Initial Release
- Chat Completion API — OpenAI-compatible endpoint with streaming, tool calls, and function calling
- Speech to Text —
saaras:v3with 22 Indic language support - Text to Speech —
bulbul:v3(37 voices across 11 Indian languages) - WhatsApp Bot — Full WhatsApp Business API integration
- Voice Calling — Twilio-based inbound voice with WebSocket streaming
- Multi-tenant — Complete tenant isolation with role-based access
- API Keys — Scoped API keys with usage tracking
- Webhook Delivery — Outbound webhooks with retry and HMAC signing
- Analytics Dashboard — Real-time conversation and usage analytics
- Model catalog — LLM, STT, TTS and image models from one OpenAI-compatible endpoint