Skip to main content

Bring Your Own LLM

Run a voice agent's replies on your own OpenAI-compatible chat completions endpoint, while CallMissed handles the call, listening and speech.

Overview

A voice agent normally answers with the model in its voice_model setting. With bring your own LLM you register your own OpenAI-compatible chat completions endpoint, point an agent at it with custom_llm_id, and every reply on that agent's calls is generated by your model. CallMissed still runs the call: speech recognition, turn-taking, the voice, tools, transcripts and phone carriage.

Your endpoint receives the conversation as a standard streamed chat completions request, tool definitions included, and streams its reply back. Your model's tokens are yours: CallMissed does not bill them. A call that runs on your own LLM pays a per-minute platform fee in their place; speech recognition, voice and phone minutes bill as they do on any call.

Bring your own LLM is not yet available on every account. Until it is on yours, creating an endpoint, testing one and setting custom_llm_id return an error saying the feature is not available on this account yet.

Authentication

Every endpoint on this page takes Authorization: Bearer cm_your_api_key. Reading endpoints needs the bots:read scope; creating, changing, deleting and testing needs bots:write. A key without the scope gets 403.

The request we send

For each reply, CallMissed sends one POST to {base_url}/chat/completions:

POST https://llm.example.com/v1/chat/completions
Content-Type: application/json
Accept: text/event-stream, application/json
Authorization: Bearer your-model-key

{
  "model": "your-model-id",
  "stream": true,
  "stream_options": { "include_usage": true },
  "messages": [
    { "role": "system", "content": "You are the front desk for..." },
    { "role": "user", "content": "Hi, I want to book a table for two." }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "check_availability",
        "description": "...",
        "parameters": { "type": "object", "properties": { "...": {} } }
      }
    }
  ]
}
  • model is always the model you registered, whatever the agent sends.
  • The auth header is the header_name and header_value you registered (Authorization by default). It is the only credential your endpoint sees.
  • tools is present when the agent has tools. Answer a tool call the OpenAI way (delta.tool_calls); the agent runs the tool and sends you the result on the next request.
  • temperature is sent when the agent sets llm_temperature.

Your endpoint must reply 200 with Content-Type: text/event-stream and Server-Sent Events chat completion chunks, ending with data: [DONE]:

data: {"id":"c1","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":"Sure, "}}]}

data: {"id":"c1","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"for what time?"}}]}

data: [DONE]

Keep replies short and conversational: they are spoken aloud on a live call, and the caller hears nothing until your first chunk arrives.

Timeouts and failures

timeout_seconds (1 to 10, default 10) is how long we wait for your endpoint to connect and for each chunk of the reply, the first one included. A reply that stops streaming for longer is cut off there. A single reply is also capped in total length and duration.

Any non-200 status, a timeout, a redirect or a connection failure counts as a failed reply. What happens next depends on platform_fallback:

platform_fallbackA failed reply
false (default)The request is retried a few times; if it keeps failing, the agent cannot answer that turn. Your model is never replaced
trueNo retry: the agent's own voice_model answers that turn instead, and later turns while your endpoint stays down

The status and body your endpoint returned are not passed back to the caller or stored. Use POST /api/v1/custom-llms/{id}/test to check an endpoint.

Address rules

base_url must be https:// with a public host name or address. It cannot carry a username, password, query string or fragment; a pasted full .../chat/completions URL is trimmed to its base. A host that resolves to a private, loopback, link-local or cloud metadata address is refused when you save it, and the address is checked again on every request, so a DNS change cannot point the agent at a private network. Redirects are not followed.

Endpoints

Fields

FieldTypeDefaultDescription
namestringrequired1 to 120 characters, for you
base_urlstringrequiredYour OpenAI-compatible base URL, for example https://llm.example.com/v1. See address rules above
modelstringrequiredThe model id sent to your endpoint, up to 200 characters
bot_iduuidnullOnly this agent may use the endpoint. null lets any of your agents use it. Set at creation
header_namestringAuthorizationThe header your endpoint authenticates with. Letters, digits and hyphens; headers such as Host, Content-Type or Cookie are refused
header_valuestringnullThe header's value, up to 4096 printable characters. Write-only: stored encrypted and never returned
timeout_secondsint101 to 10, see above
platform_fallbackboolfalseAnswer a failed reply with the agent's voice_model
enabledbooltrueA disabled endpoint fails every reply (or falls back, when platform_fallback is on)

Responses carry key_last4 (the last four characters of header_value, or null when the value is shorter than 12 characters) in place of the value, plus id, created_at and updated_at. An account can register up to 50 endpoints.

Register an endpoint

curl -X POST https://api.callmissed.com/api/v1/custom-llms \
  -H "Authorization: Bearer cm_your_api_key" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Support model",
    "base_url": "https://llm.example.com/v1",
    "model": "support-v3",
    "header_name": "Authorization",
    "header_value": "Bearer your-model-key",
    "timeout_seconds": 8,
    "platform_fallback": true
  }'

201 returns the endpoint:

{
  "id": "6f1c2a0e-6c8f-4a51-9d0e-2f7e1b9c4d11",
  "bot_id": null,
  "name": "Support model",
  "base_url": "https://llm.example.com/v1",
  "model": "support-v3",
  "header_name": "Authorization",
  "key_last4": "-key",
  "timeout_seconds": 8,
  "platform_fallback": true,
  "enabled": true,
  "created_at": "2026-10-08T09:30:00Z",
  "updated_at": "2026-10-08T09:30:00Z"
}
StatusMeaning
403The key lacks bots:write, or the feature is not available on this account yet
404bot_id is not one of your agents
409The account already has 50 endpoints
422A field is invalid, or base_url is not https or does not resolve to a public address

List, get, update, delete

# List (optional ?bot_id=, limit 1-100, offset)
curl https://api.callmissed.com/api/v1/custom-llms \
  -H "Authorization: Bearer cm_your_api_key"

# Get one
curl https://api.callmissed.com/api/v1/custom-llms/{id} \
  -H "Authorization: Bearer cm_your_api_key"

# Update: send only what changes. "header_value": null removes the stored header.
curl -X PATCH https://api.callmissed.com/api/v1/custom-llms/{id} \
  -H "Authorization: Bearer cm_your_api_key" \
  -H "Content-Type: application/json" \
  -d '{"model": "support-v4", "header_value": "Bearer rotated-key"}'

# Delete
curl -X DELETE https://api.callmissed.com/api/v1/custom-llms/{id} \
  -H "Authorization: Bearer cm_your_api_key"

A delete returns 204, or 409 while an agent's custom_llm_id still points at the endpoint; the message names those agents. An unknown id, or one in another account, is 404.

Test an endpoint

curl -X POST https://api.callmissed.com/api/v1/custom-llms/{id}/test \
  -H "Authorization: Bearer cm_your_api_key"

Sends one short streamed request ("Reply with the single word OK.") exactly as a call would, and reports whether a well-formed chat completion chunk came back. The reply's text is not returned or stored.

{ "ok": true, "status_code": 200, "latency_ms": 412, "error": null }

On failure ok is false, status_code is your endpoint's status (or null when it could not be reached) and error says what went wrong in general terms, for example "The endpoint returned HTTP 401".

Point an agent at it

Set custom_llm_id in the agent's config:

curl -X PATCH https://api.callmissed.com/api/v1/bots/{bot_id}/config \
  -H "Authorization: Bearer cm_your_api_key" \
  -H "Content-Type: application/json" \
  -d '{"values": {"custom_llm_id": "6f1c2a0e-6c8f-4a51-9d0e-2f7e1b9c4d11"}}'

To go back to the agent's own model, unset it: {"unset": ["custom_llm_id"]}.

custom_llm_id cannot be combined with:

  • a voice_tier: a tier runs its own model at a flat per-minute price;
  • a speech-to-speech voice_model (such as the realtime or managed voice models), which listens and speaks itself and so has no separate reply step.

Either combination is refused with 422 and a message saying which setting to change. The agent's voice_model stays relevant as the fallback model when platform_fallback is on.

Each call starts with your endpoint's current settings, so a changed model, header or timeout applies from the next call. An endpoint pinned to one agent with bot_id cannot serve any other agent's calls.