Skip to main content

Managed Voice Agent

Stream audio over a WebSocket and get a speaking agent back. Two protocols: CallMissed-native and Deepgram Voice Agent compatible.

Overview

The Managed Voice Agent is a full speech-to-speech pipeline behind a single WebSocket. You stream microphone audio in; you get synthesized speech and conversation events back. Speech recognition, the language model, text-to-speech, turn-taking and interruption handling are all run and tuned for you.

Unlike the Voice Session API, there is no WebRTC and no client SDK to install — a plain WebSocket and raw PCM is the whole integration.

Two wire protocols are served, backed by the same engine:

EndpointProtocol
wss://api.callmissed.com/v2/voice/agentCallMissed-native
wss://api.callmissed.com/v1/agent/converseDeepgram Voice Agent compatible

If you already have an integration written against Deepgram's Voice Agent API, point it at the second URL and it will work unchanged.

Both live on api.callmissed.com, the same host as the rest of the API — there is no separate hostname to allowlist.

Authentication

Both endpoints take an API key with stt, tts and llm permissions.

Authorization: Token cm_your_api_key

Authorization: Bearer cm_your_api_key is accepted too.

Browsers cannot set headers on a WebSocket, so the subprotocol form is also accepted:

new WebSocket(url, ["token", "cm_your_api_key"])

A missing or invalid key, a key without those permissions, an exhausted plan limit, the concurrent-session cap and a credit balance below the minimum are all checked before the socket is accepted, so a refused connection fails at the handshake rather than mid-conversation.

Session flow

  1. Connect. The server sends Welcome with a request_id.
  2. Send Settings as your first message, within 15 seconds. Anything else first (including audio) ends the session with an Error.
  3. Wait for SettingsApplied, then start streaming audio.
  4. Send raw PCM as binary frames; receive synthesized PCM as binary frames and events as JSON on the same socket.
{ "type": "Welcome", "request_id": "fc553ec9-5874-49ca-a47c-b670d525a4b1" }

Settings (native)

The native shape is organised around what you actually choose: one model id per role.

{
  "type": "Settings",
  "audio": {
    "input":  { "encoding": "linear16", "sample_rate": 24000 },
    "output": { "encoding": "linear16", "sample_rate": 24000 }
  },
  "agent": {
    "prompt": "You are a concise support agent for an Indian retail brand.",
    "greeting": "Hi, how can I help?",
    "language": "en-IN",
    "llm": { "model": "gpt-oss-120b", "temperature": 0.4 },
    "stt": { "model": "saaras:v3" },
    "tts": { "model": "bulbul:v3", "voice": "shubh" },
    "tools": [
      {
        "name": "lookup_order",
        "description": "Look up an order by id",
        "parameters": {
          "type": "object",
          "properties": { "order_id": { "type": "string" } },
          "required": ["order_id"]
        }
      }
    ]
  },
  "tags": ["support"]
}
FieldDefaultNotes
audio.input / audio.outputlinear16 at 24000 HzSee Audio format
agent.prompta concise built-in assistant promptString. Longer than 25,000 characters is truncated, with a PROMPT_TOO_LONG warning
agent.greeting—The first line the agent speaks
agent.languageen-INBCP-47. Falls back to agent.stt.language, then agent.tts.language
agent.llm.modelgemma-4-31bAny id from GET /api/v1/voice/models
agent.llm.temperaturemodel defaultNumber, 0–2
agent.stt.modelsaaras:v3
agent.tts.modelplan-dependentFree plan bulbul:v3; paid plans (starter, pro, enterprise) sonic-3.6
agent.tts.voiceplan-dependentFree plan shubh, paid plans skylar. Must be a speaker of the chosen model, so set it whenever you set agent.tts.model
agent.tools[]Functions the agent may call: name (required), description, parameters (JSON Schema)
tags[]Array of strings

agent.llm, agent.stt and agent.tts also accept a bare model id string as shorthand, e.g. "llm": "kimi-k2.5".

A model id the voice agent cannot serve fails Settings with an INVALID_SETTINGS error naming it; it is never swapped for another model.

Settings (Deepgram-compatible)

On /v1/agent/converse, send Deepgram's Voice Agent Settings shape unchanged, with CallMissed model ids in agent.listen.provider.model, agent.think.provider.model and agent.speak.provider.model (model_id is accepted too). agent.think.prompt, agent.think.functions, agent.greeting, the audio block and tags behave as on the native surface.

Differences from Deepgram, each reported rather than silently dropped:

  • A fallback array in agent.think or agent.speak uses only its first entry (THINK_FALLBACK_IGNORED / SPEAK_FALLBACK_IGNORED warning).
  • Fields accepted but not applied (for example listen.provider.keyterms, speak.provider.speed, think.provider.credentials) are listed in one SETTINGS_FIELDS_IGNORED warning.
  • audio.output.container must be none (raw frames).
  • A function with a server-side endpoint is dispatched to your client instead (see Tool calling).

Every other client message — updates, injection, tool responses, keepalive — is identical across both protocols. Switching protocols means rewriting one message. Settings is applied once; sending it again returns a DUPLICATE_SETTINGS warning.

Audio format

linear16 (raw 16-bit PCM, little-endian, mono) in both directions, at 8000, 16000, 24000, 32000 or 48000 Hz. Output is raw frames with no container: a WAV or OGG header would be read as audio by most telephony consumers.

An unsupported encoding, sample rate or container is rejected at Settings time rather than accepted and quietly changed, so you find out at the handshake instead of hearing the wrong thing on a call.

Choosing models

Call GET /api/v1/voice/models for the models you can use, each with its measured latency. Settings accepts any id in that response; filter on eligible to pick models measured fast enough to hold a conversation. Any speech-to-text × language model × text-to-speech combination is valid.

Eligibility is based on measured performance from real traffic, not on vendor claims. A model that is too slow to hold a conversation is marked ineligible and listed with the reason, so you can see why rather than wondering where it went.

Two speech-to-text models are streaming-capable here that you cannot use for file transcription in the same way:

  • ink-2 ($0.5625 / hr) — Cartesia's top-ranked voice-agent STT: 8% WER on AppTek's 14-accent call-centre benchmark, vs 10% Deepgram Flux and 12% ElevenLabs. It self-detects turns. English only — set "language": "en". Sending non-English audio does not error; it just transcribes badly. This is the only surface ink-2 runs on: the file transcription endpoint rejects it with a 400.
  • ink-whisper ($0.1875 / hr) — 100 languages including Hindi, Urdu and Tamil. Cartesia's cheapest STT, with dynamic chunking that reduces hallucination across pauses. Use this instead of ink-2 for any non-English call.

List available models

curl https://api.callmissed.com/api/v1/voice/models \
  -H "Authorization: Bearer cm_your_api_key"
{
  "turn_budget_ms": 1000,
  "llm": [
    {
      "id": "gpt-oss-120b",
      "label": "GPT-OSS 120B",
      "eligible": true,
      "verdict": "eligible",
      "p50_ms": 180.0,
      "samples": 240,
      "budget_ms": 400,
      "reason": "p50 180ms within the 400ms budget",
      "languages": ["en"],
      "voices": []
    }
  ],
  "stt": [],
  "tts": []
}
verdictMeaning
eligibleMeasured within its stage budget. Selectable.
too_slowMeasured over budget. Not offered; reason says by how much.
unsupportedStructurally unavailable (e.g. under maintenance).
unmeasuredNot enough measured turns yet to give a verdict.

p50_ms is null when a model has no measurement. It is never reported as 0 — zero would read as "instant" for a stage that was simply never measured.

FieldMeaning
turn_budget_msThe voice-to-voice target the three stage budgets add up to
id / labelModel id to send, and its display name
eligibletrue when verdict is eligible — the field to filter on
p50_ms / samplesMedian measured latency for the model's stage over the last 30 days, and the number of measured turns behind it
budget_msThe stage budget the model is judged against
reasonWhy the model got its verdict
languagesBCP-47 codes where recorded; [] means not recorded
voicesSelectable speakers (TTS models only; [] otherwise)

Server events

EventFieldsMeaning
Welcomerequest_idSocket open.
SettingsApplied—Configuration accepted; start streaming.
ConversationTextrole (user / assistant), contentA finished turn.
UserStartedSpeaking—Barge-in. Stop playback and clear your buffer.
AgentThinkingcontentThe model is working.
AgentStartedSpeakingtotal_latency, tts_latency, ttt_latencyFirst audio of a turn. Seconds.
AgentAudioDone—Last audio chunk sent for this turn.
FunctionCallRequestfunctions[]Run a tool and reply.
LatencyReporteou_latency, stt_latency, ttt_latency, tts_latency, total_latencyPer-stage timings for the turn. Seconds.
PromptUpdated / ThinkUpdated / ListenUpdated / SpeakUpdated—The matching update was applied.
InjectionRefused—An InjectAgentMessage was refused (see below).
Warningcode, descriptionNon-fatal; the session continues.
Errorcode, descriptionFatal; the socket closes. Reconnect to start a new session.

UserStartedSpeaking is the only barge-in signal — there is no separate flush message. When you receive it, stop playback and discard whatever audio you have buffered. If you don't, the caller keeps hearing the interrupted sentence for as long as your playback buffer is deep, which is the most common reason interruption appears not to work.

AgentAudioDone means the last chunk was sent, not that the caller heard it. Your own playback buffer may still be draining.

Latency fields are omitted when a stage was not measured, rather than reported as 0.

Error codes

codeCause
SETTINGS_TIMEOUTNo Settings within 15 seconds of Welcome
NON_SETTINGS_MESSAGE_BEFORE_SETTINGSThe first message was not Settings
INVALID_JSONSettings was not valid JSON
INVALID_SETTINGSA field failed validation, or a model id cannot be served
INSUFFICIENT_CREDITSYour balance ran out mid-session
SESSION_TIMEOUTThe 2-hour session cap was reached
AGENT_ERROR / INTERNAL_ERRORThe voice engine could not continue or could not start

Common Warning codes: INVALID_UPDATE, THINK_UPDATE_FAILED, LISTEN_UPDATE_FAILED, SPEAK_UPDATE_FAILED, FUNCTION_CALL_TIMEOUT, UNMATCHED_FUNCTION_RESPONSE, FUNCTION_ENDPOINT_UNSUPPORTED, UNKNOWN_MESSAGE, DUPLICATE_SETTINGS, INVALID_JSON, MESSAGE_TOO_LARGE, AUDIO_CHUNK_TOO_LARGE (the frame is dropped), PROMPT_TOO_LONG.

Tool calling

Declare tools in Settings (agent.tools on the native surface, agent.think.functions on the Deepgram-compatible one), then answer requests over the same socket.

{
  "type": "FunctionCallRequest",
  "functions": [
    {
      "id": "fc_01H...",
      "name": "lookup_order",
      "arguments": "{\"order_id\":\"A-1042\"}",
      "client_side": true
    }
  ]
}

Reply with the result. arguments is a JSON string, not an object. Echo the id; content should be a string (anything else is JSON-encoded for you).

{
  "type": "FunctionCallResponse",
  "id": "fc_01H...",
  "name": "lookup_order",
  "content": "{\"status\":\"shipped\",\"eta\":\"2 days\"}"
}

A tool that does not answer within 30 seconds does not hang the call — the agent tells the caller it could not complete the action, and a FUNCTION_CALL_TIMEOUT warning is emitted.

Server-side execution (a functions[].endpoint) is not supported. Declaring one returns a Warning and the call is dispatched to your client instead, so an unsupported mode is never silently ignored.

Updating a live session

MessageBodyEffectAcknowledged by
UpdatePromptpromptReplaces the system promptPromptUpdated
UpdateThinkmodel, optional temperature, promptSwitch the language model (and optionally the prompt)ThinkUpdated
UpdateListenmodel, optional languageSwitch speech recognitionListenUpdated
UpdateSpeakmodel, voice, optional languageSwitch voice or text-to-speech modelSpeakUpdated
InjectAgentMessagemessage, optional behaviorMake the agent say something nowInjectionRefused if refused
InjectUserMessagecontentInject text as if the caller said it—
KeepAlive—Hold an idle socket open (does not extend the 2-hour cap)—
{ "type": "UpdateSpeak", "model": "sonic-3.6", "voice": "skylar" }

The update bodies can be sent flat as above, or nested the Deepgram way ({"type": "UpdateSpeak", "speak": {"provider": {"model": "sonic-3.6"}}}) on either surface.

InjectAgentMessage behavior is default, queue or interrupt. default is refused while either side is speaking, queue only while the caller is speaking, and interrupt is never refused.

A model id the service cannot serve keeps the current model and returns a Warning; it is never swapped for a different one behind your back.

Latency

The fast path targets sub-1s from end of your speech to first audio back. That budget is the sum of three serial stages:

StageBudget
Speech recognition, final transcript300 ms
Language model, first token400 ms
Text-to-speech, first byte300 ms

LatencyReport gives you the real numbers per turn, so you can measure rather than take our word for it.

Sub-200ms end-to-end voice-to-voice is not achievable with a speech-to-text → language model → speech pipeline, by anyone. Detecting that you stopped speaking alone costs more than that. Treat sub-second as the realistic target and measure the rest with LatencyReport.

Limits

LimitValue
Maximum session length2 hours
Time to send Settings15 seconds
Tool response timeout30 seconds
Concurrent sessionsPer plan

Usage is billed per turn across speech recognition, the language model and text-to-speech, at each model's own rate. If your balance runs out mid-session you receive an INSUFFICIENT_CREDITS Error frame and the socket closes, rather than the call continuing unbilled.