Bring Your Own LLM
Run a voice agent's replies on your own OpenAI-compatible chat completions endpoint, while CallMissed handles the call, listening and speech.
Overview
A voice agent normally answers with the model in its voice_model setting. With bring your own LLM you register your own OpenAI-compatible chat completions endpoint, point an agent at it with custom_llm_id, and every reply on that agent's calls is generated by your model. CallMissed still runs the call: speech recognition, turn-taking, the voice, tools, transcripts and phone carriage.
Your endpoint receives the conversation as a standard streamed chat completions request, tool definitions included, and streams its reply back. Your model's tokens are yours: CallMissed does not bill them. A call that runs on your own LLM pays a per-minute platform fee in their place; speech recognition, voice and phone minutes bill as they do on any call.
Bring your own LLM is not yet available on every account. Until it is on yours, creating an endpoint, testing one and setting custom_llm_id return an error saying the feature is not available on this account yet.
Authentication
Every endpoint on this page takes Authorization: Bearer cm_your_api_key. Reading endpoints needs the bots:read scope; creating, changing, deleting and testing needs bots:write. A key without the scope gets 403.
The request we send
For each reply, CallMissed sends one POST to {base_url}/chat/completions:
POST https://llm.example.com/v1/chat/completions
Content-Type: application/json
Accept: text/event-stream, application/json
Authorization: Bearer your-model-key
{
"model": "your-model-id",
"stream": true,
"stream_options": { "include_usage": true },
"messages": [
{ "role": "system", "content": "You are the front desk for..." },
{ "role": "user", "content": "Hi, I want to book a table for two." }
],
"tools": [
{
"type": "function",
"function": {
"name": "check_availability",
"description": "...",
"parameters": { "type": "object", "properties": { "...": {} } }
}
}
]
}modelis always themodelyou registered, whatever the agent sends.- The auth header is the
header_nameandheader_valueyou registered (Authorizationby default). It is the only credential your endpoint sees. toolsis present when the agent has tools. Answer a tool call the OpenAI way (delta.tool_calls); the agent runs the tool and sends you the result on the next request.temperatureis sent when the agent setsllm_temperature.
Your endpoint must reply 200 with Content-Type: text/event-stream and Server-Sent Events chat completion chunks, ending with data: [DONE]:
data: {"id":"c1","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":"Sure, "}}]}
data: {"id":"c1","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"for what time?"}}]}
data: [DONE]Keep replies short and conversational: they are spoken aloud on a live call, and the caller hears nothing until your first chunk arrives.
Timeouts and failures
timeout_seconds (1 to 10, default 10) is how long we wait for your endpoint to connect and for each chunk of the reply, the first one included. A reply that stops streaming for longer is cut off there. A single reply is also capped in total length and duration.
Any non-200 status, a timeout, a redirect or a connection failure counts as a failed reply. What happens next depends on platform_fallback:
platform_fallback | A failed reply |
|---|---|
false (default) | The request is retried a few times; if it keeps failing, the agent cannot answer that turn. Your model is never replaced |
true | No retry: the agent's own voice_model answers that turn instead, and later turns while your endpoint stays down |
The status and body your endpoint returned are not passed back to the caller or stored. Use POST /api/v1/custom-llms/{id}/test to check an endpoint.
Address rules
base_url must be https:// with a public host name or address. It cannot carry a username, password, query string or fragment; a pasted full .../chat/completions URL is trimmed to its base. A host that resolves to a private, loopback, link-local or cloud metadata address is refused when you save it, and the address is checked again on every request, so a DNS change cannot point the agent at a private network. Redirects are not followed.
Endpoints
Fields
| Field | Type | Default | Description |
|---|---|---|---|
name | string | required | 1 to 120 characters, for you |
base_url | string | required | Your OpenAI-compatible base URL, for example https://llm.example.com/v1. See address rules above |
model | string | required | The model id sent to your endpoint, up to 200 characters |
bot_id | uuid | null | Only this agent may use the endpoint. null lets any of your agents use it. Set at creation |
header_name | string | Authorization | The header your endpoint authenticates with. Letters, digits and hyphens; headers such as Host, Content-Type or Cookie are refused |
header_value | string | null | The header's value, up to 4096 printable characters. Write-only: stored encrypted and never returned |
timeout_seconds | int | 10 | 1 to 10, see above |
platform_fallback | bool | false | Answer a failed reply with the agent's voice_model |
enabled | bool | true | A disabled endpoint fails every reply (or falls back, when platform_fallback is on) |
Responses carry key_last4 (the last four characters of header_value, or null when the value is shorter than 12 characters) in place of the value, plus id, created_at and updated_at. An account can register up to 50 endpoints.
Register an endpoint
curl -X POST https://api.callmissed.com/api/v1/custom-llms \
-H "Authorization: Bearer cm_your_api_key" \
-H "Content-Type: application/json" \
-d '{
"name": "Support model",
"base_url": "https://llm.example.com/v1",
"model": "support-v3",
"header_name": "Authorization",
"header_value": "Bearer your-model-key",
"timeout_seconds": 8,
"platform_fallback": true
}'201 returns the endpoint:
{
"id": "6f1c2a0e-6c8f-4a51-9d0e-2f7e1b9c4d11",
"bot_id": null,
"name": "Support model",
"base_url": "https://llm.example.com/v1",
"model": "support-v3",
"header_name": "Authorization",
"key_last4": "-key",
"timeout_seconds": 8,
"platform_fallback": true,
"enabled": true,
"created_at": "2026-10-08T09:30:00Z",
"updated_at": "2026-10-08T09:30:00Z"
}| Status | Meaning |
|---|---|
403 | The key lacks bots:write, or the feature is not available on this account yet |
404 | bot_id is not one of your agents |
409 | The account already has 50 endpoints |
422 | A field is invalid, or base_url is not https or does not resolve to a public address |
List, get, update, delete
# List (optional ?bot_id=, limit 1-100, offset)
curl https://api.callmissed.com/api/v1/custom-llms \
-H "Authorization: Bearer cm_your_api_key"
# Get one
curl https://api.callmissed.com/api/v1/custom-llms/{id} \
-H "Authorization: Bearer cm_your_api_key"
# Update: send only what changes. "header_value": null removes the stored header.
curl -X PATCH https://api.callmissed.com/api/v1/custom-llms/{id} \
-H "Authorization: Bearer cm_your_api_key" \
-H "Content-Type: application/json" \
-d '{"model": "support-v4", "header_value": "Bearer rotated-key"}'
# Delete
curl -X DELETE https://api.callmissed.com/api/v1/custom-llms/{id} \
-H "Authorization: Bearer cm_your_api_key"A delete returns 204, or 409 while an agent's custom_llm_id still points at the endpoint; the message names those agents. An unknown id, or one in another account, is 404.
Test an endpoint
curl -X POST https://api.callmissed.com/api/v1/custom-llms/{id}/test \
-H "Authorization: Bearer cm_your_api_key"Sends one short streamed request ("Reply with the single word OK.") exactly as a call would, and reports whether a well-formed chat completion chunk came back. The reply's text is not returned or stored.
{ "ok": true, "status_code": 200, "latency_ms": 412, "error": null }On failure ok is false, status_code is your endpoint's status (or null when it could not be reached) and error says what went wrong in general terms, for example "The endpoint returned HTTP 401".
Point an agent at it
Set custom_llm_id in the agent's config:
curl -X PATCH https://api.callmissed.com/api/v1/bots/{bot_id}/config \
-H "Authorization: Bearer cm_your_api_key" \
-H "Content-Type: application/json" \
-d '{"values": {"custom_llm_id": "6f1c2a0e-6c8f-4a51-9d0e-2f7e1b9c4d11"}}'To go back to the agent's own model, unset it: {"unset": ["custom_llm_id"]}.
custom_llm_id cannot be combined with:
- a
voice_tier: a tier runs its own model at a flat per-minute price; - a speech-to-speech
voice_model(such as the realtime or managed voice models), which listens and speaks itself and so has no separate reply step.
Either combination is refused with 422 and a message saying which setting to change. The agent's voice_model stays relevant as the fallback model when platform_fallback is on.
Each call starts with your endpoint's current settings, so a changed model, header or timeout applies from the next call. An endpoint pinned to one agent with bot_id cannot serve any other agent's calls.
Voice Agent Tools
What a voice agent can call mid-conversation on a phone, WhatsApp or WebRTC call: built-in tools, your own REST tools, MCP servers, and the two tools that are always there.
Agent Evals
Regression-test a voice agent against scripted personas with pass/fail assertions, and read the transcript of every case.