Skip to main content

Real-time STT

Real-time speech-to-text transcription via WebSocket.

Overview

Real-time STT is available through the Voice Agent WebSocket pipeline. Audio is streamed as raw PCM s16le 16 kHz mono, and a transcript is returned for each utterance as the user speaks.

There is no standalone real-time STT WebSocket endpoint — real-time transcription is part of the full Voice Agent pipeline (STT → LLM → TTS).

For file-based transcription, use the Speech to Text REST API.

Via Voice Agent

Connect to WS /ws/voice-agent, send a config message, then stream raw audio. Each finished user utterance comes back as a transcript message:

{"type": "transcript", "text": "Hello, how are you?"}

The socket also runs the agent's reply (LLM + speech), so you receive llm_response, agent_text, audio_start / audio_end and binary MP3 audio as well. A transcript-only client can ignore those.

Authentication

Pass your API key as a WebSocket subprotocol: Sec-WebSocket-Protocol: token, cm_your_key. That is a request header, so the key stays out of access logs and proxy history, which a query string does not. In the browser the constructor's second argument sets it, and the order matters: the literal token first, then the key.

new WebSocket(url, ["token", "cm_your_key"]);

Clients that can set headers may send Authorization: Bearer cm_your_key instead.

The ?key=cm_your_key query parameter is deprecated. It still works so existing integrations keep connecting, but prefer the subprotocol for anything new.

The key needs the stt, tts and llm permissions, because the socket runs all three.

Audio format

Send audio as binary frames of raw PCM: signed 16-bit little-endian, 16 kHz, mono, with no container header. Keep frames small (tens to hundreds of milliseconds of audio); oversized frames are dropped. Compressed audio (WebM/Opus from MediaRecorder, MP3) is not decoded.

Config message

The first text frame must be the config, sent within 10 seconds of connecting. Every field is optional. Send another config later to change settings mid-session; the server answers each one with {"type": "ready"}.

FieldTypeDescription
typestring"config"
languagestringSpeech locale, default en-IN. One of bn-IN, en-IN, gu-IN, hi-IN, kn-IN, ml-IN, mr-IN, od-IN, pa-IN, ta-IN, te-IN (a bare hi is read as hi-IN); anything else falls back to en-IN
voicestringA bulbul:v3 voice from Voices, default shubh. An unknown voice falls back to the language's default voice
system_promptstringInstructions for the agent's replies
llm_modelstringChat model ID for the replies, default gemma-4-31b. Must be a chat model available on your plan

Send {"type": "clear_history"} to reset the conversation memory.

Close codes

CodeMeaning
4001Missing or invalid API key
4003Key lacks stt/tts/llm permission, or llm_model is not allowed
4008Plan usage limit reached, or insufficient credits
4009Config message too large
1013Too many connections — retry later

Example

const ws = new WebSocket(
  "wss://api.callmissed.com/ws/voice-agent",
  ["token", "cm_your_key"]
);
ws.binaryType = "arraybuffer";

ws.onopen = async () => {
  ws.send(JSON.stringify({ type: "config", language: "hi-IN", voice: "shubh" }));

  // Capture the microphone at 16 kHz and send raw 16-bit PCM.
  const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
  const ctx = new AudioContext({ sampleRate: 16000 });
  const source = ctx.createMediaStreamSource(stream);
  const processor = ctx.createScriptProcessor(4096, 1, 1);
  processor.onaudioprocess = (e) => {
    const f32 = e.inputBuffer.getChannelData(0);
    const i16 = new Int16Array(f32.length);
    for (let i = 0; i < f32.length; i++) {
      i16[i] = Math.max(-1, Math.min(1, f32[i])) * 0x7fff;
    }
    ws.send(i16.buffer);
  };
  source.connect(processor);
  processor.connect(ctx.destination);
};

ws.onmessage = (event) => {
  if (typeof event.data === "string") {
    const msg = JSON.parse(event.data);
    if (msg.type === "transcript") console.log("User said:", msg.text);
    if (msg.type === "error") console.error(msg.code, msg.message);
  }
  // Binary frames are the agent's MP3 reply audio.
};

See the Voice Agent page for the full WebSocket protocol and all message types.