Skip to main content

Chat Completion

Generate text responses using our OpenAI-compatible chat completion API.

Overview

The Chat Completion API generates AI responses given a list of messages. It's fully OpenAI-compatible — use the same SDK and request format.

Endpoint: POST /v1/chat/completions

How a request flows

Every chat completion takes the same path through the platform — your app never talks to the underlying provider directly:

  1. 1

    Your app

    Send with +

  2. 2

    CallMissed gateway

    Authenticate the key, check credits, route by model id

  3. 3

    Provider

    Run inference on the best-fit backend — picked from the model id

  4. 4

    CallMissed gateway

    Stream tokens back and deduct credits when the response completes

  5. 5

    Your app

    Receive the completion (all at once, or token-by-token when streaming)

Tip: The model id decides routing automatically — you never pick a backend. See How CallMissed Works.

Make your first request

Get an API key

Create a key in the dashboard (Developer → API keys). It looks like cm_xxxx… and is shown once.

Point your SDK at CallMissed

Set the base URL to https://api.callmissed.com/v1 and pass your cm_ key. No other change to your OpenAI code.

Send messages and read the reply

Call chat.completions.create with a model and a messages array. Read response.choices[0].message.content.

Basic completion

Send a single-turn or multi-turn conversation and receive a complete response. Use any OpenAI SDK — set base_url to https://api.callmissed.com/v1 and api_key to your cm_ key.

Streaming

Set stream: true to receive tokens as they're generated. See Streaming for full examples.

Function calling

Pass a tools array to let the model call your functions. See Function Calling.

Basic Usage

from openai import OpenAI

client = OpenAI(
    api_key="cm_your_key",
    base_url="https://api.callmissed.com/v1"
)

response = client.chat.completions.create(
    model="sarvam-105b",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "What is the capital of India?"}
    ]
)

print(response.choices[0].message.content)

Parameters

ParameterTypeDescription
modelstringModel ID (e.g. sarvam-105b, gpt-5.6-luna). Defaults to sarvam-105b when omitted
messagesarrayRequired. 1–2,000 {role, content} objects. System prompt goes here as {"role": "system", "content": "..."}. content may be a string or an array of text, image, audio or document parts (see Vision and Audio and documents)
streambooleanEnable streaming SSE responses (default false)
stream_optionsobject{"include_usage": true} to get token counts in the stream
temperaturenumberSampling temperature, 0–2
max_tokensintegerMaximum tokens to generate, 1–1,048,576. A model with a documented output limit (the max_output_tokens field in GET /api/v1/models) rejects a larger value with 400 max_tokens_too_large
max_completion_tokensintegerSame as max_tokens (the newer OpenAI name); used when max_tokens is absent
nintegerNumber of completions, 1–5 (default 1)
top_pnumberNucleus sampling, 0–1
top_kintegerTop-K sampling, 1–200
frequency_penaltynumberPenalize repeated tokens, -2–2
presence_penaltynumberPenalize new topics, -2–2
repetition_penaltynumberReduce repetition, 0–2
seedintegerDeterministic sampling, 0–2^63-1
stopstring or arrayUp to 16 stop sequences
logit_biasobjectToken id → bias between -100 and 100; at most 1,024 entries
logprobsbooleanReturn log probabilities
top_logprobsintegerTop N log probs per token, 0–20
toolsarrayUp to 128 function definitions — see Function Calling
tool_choicestring or object"auto", "none", "required", or {"type": "function", "function": {"name": "..."}}
parallel_tool_callsbooleanAllow parallel function calls
response_formatobject{"type": "json_object"} or {"type": "json_schema", "json_schema": {...}}
structured_outputsbooleanEnforce strict JSON schema
reasoning_effortstring"none" / "minimal" / "low" / "medium" / "high" / "xhigh" — each model accepts a different subset and the gateway maps the rest; see the per-model matrix
modelsarrayUp to 32 fallback model ids, tried in order if model fails with a retryable error — see Model substitution
userstringYour end user's id. Enforces that user's monthly budget when one is set (caps apply to ids of at most 256 characters)
providerobject{"zdr": true} serves the request only on a zero-data-retention route, or fails with zdr_unavailable
prompt_cache_keystringUp to 1,024 characters. Reuse the same key for requests that share a long prompt prefix to raise the cache hit rate — see Prompt caching
bot_idstringID of one of your bots. Its knowledge base is searched with the latest user message and the top matches are added to the prompt — see Knowledge
knowledge_top_kintegerWith bot_id: how many knowledge chunks to add, 1–50 (default 6)
knowledge_min_scorenumberWith bot_id: minimum similarity score, 0–1 (default 0)

On GPT-5 and GPT-6 family models, temperature, top_p and logit_bias are not supported by the model and are left out of the request; the model runs at its default sampling.

OpenAI Python SDK note — The OpenAI client validates kwargs against its known parameters, so a CallMissed-specific field such as reasoning_effort raises TypeError: Completions.create() got an unexpected keyword argument. Pass it via extra_body instead:

client.chat.completions.create(
    model="kimi-k2.6",
    messages=[...],
    extra_body={"reasoning_effort": "none"},
)

Raw HTTP / curl users can keep it at the top level — only the OpenAI SDK gates kwargs.

Model Substitution

CallMissed never substitutes your model on its own. Send a model and you get that model, or a clean error (429/503 with Retry-After).

To opt in to failover, list fallbacks yourself in models:

{
  "model": "kimi-k2.6",
  "models": ["kimi-k2.5", "gpt-oss-120b"],
  "messages": [{"role": "user", "content": "Hello"}]
}

If model fails with a retryable error (an upstream outage or rate limit), the next id in models is tried. Each fallback must pass the same checks as model — your plan, the key's allowed_models, maintenance status, vision support and context window — and ids that fail them are skipped. The response's model field names the model that actually answered, and you are billed at that model's rate. Requests that send tools, tool_choice, response_format or structured_outputs never fall back, because a different model could change the result shape.

Need a model that is not in the catalog? See Models on demand.

Vision (Image Input)

Multimodal content (text + image parts) is accepted on any model whose supports_vision flag is true in GET /v1/models. Models without vision support reject image content with 400 unsupported_image_input before the upstream call, so you're not charged.

from openai import OpenAI

client = OpenAI(api_key="cm_your_key", base_url="https://api.callmissed.com/v1")

resp = client.chat.completions.create(
    model="gpt-5.6-sol",   # supports_vision: true
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this image?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/cat.png"}},
        ],
    }],
)

Vision-capable models: gpt-6.1-sol, gpt-6-sol, gpt-6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, gpt-5.5, gpt-4o, gpt-4.1, gpt-5-mini, grok-4.3, gemini-3.8-flash, gemini-3.7-flash, gemini-3.6-flash, gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.1-pro-preview, gemini-3.1-flash-lite, gemma-4-31b, gemma-4-26b-a4b-it, kimi-k2.5, kimi-k2.5-fast, kimi-k2.6, kimi-k2.7-code, mistral-small-3.1.

GET /v1/models is authoritative. Read supports_vision there rather than hard-coding this list.

An image_url may be an https:// URL or a base64 data URL. The Gemini models and gemma-4-31b only take image bytes, so CallMissed downloads a URL image for them: it must be publicly reachable without redirects, return an image content type, and be at most 20 MB (10 MB for gemma-4-31b). Gemini accepts PNG, JPEG, WEBP, HEIC and HEIF. If a download fails the request returns 400; send base64 to avoid the extra fetch.

Audio and Document Input

The Gemini models (gemini-3.8-flash, gemini-3.7-flash, gemini-3.6-flash, gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.1-pro-preview, gemini-3.1-flash-lite) also take audio and PDF parts:

{"type": "input_audio", "input_audio": {"data": "<base64>", "format": "wav"}}
{"type": "file", "file": {"file_data": "data:application/pdf;base64,<base64>", "filename": "report.pdf"}}

Audio format may be wav, mp3, aac, flac, ogg, opus, m4a, aiff or webm. A file part must carry the PDF inline in file_data; file_id references are not supported. Any other model rejects these parts with 400 unsupported_audio_input or 400 unsupported_file_input before the upstream call, so you're not charged. Content that can't be passed on as sent is rejected with a 400 that names the part, rather than being dropped.

Context Window

Every model in the catalog advertises a context_window (token count for the combined prompt + completion). The GET /v1/models response exposes it under two keys for cross-client compatibility:

  • context_window (OpenAI/CallMissed canonical name)
  • context_length (OpenAI SDK convention — same value)
from openai import OpenAI

client = OpenAI(api_key="cm_your_key", base_url="https://api.callmissed.com/v1")

for m in client.models.list():
    extra = m.model_extra or {}
    print(m.id, extra.get("context_window"), extra.get("supports_vision"))

Snapshot — GET /v1/models is authoritative:

Modelcontext_window
gpt-5.5, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, gpt-6-sol, gpt-6-luna, gpt-6.1-sol1,050,000
DeepSeek-V4-Pro, DeepSeek-V4-Flash, glm-5.31,048,576
gpt-4.1300,000
gpt-5-mini400,000
kimi-k2.6, kimi-k2.7-code, glm-5.2262,144
kimi-k2.5, kimi-k2.5-fast, nemotron-3-super, gemma-4-26b-a4b-it256,000
grok-4.3200,000
sarvam-105b, glm-4.7-flash, gemma-4-31b131,072
gpt-4o, gpt-oss-120b, mistral-small-3.1128,000
sarvam-105b-conversations32,000

Images, audio and files count toward the window by what the model reads, not by their encoded size, so a large base64 image does not by itself cause context_length_exceeded. A prompt that is clearly longer than the window is rejected with 400 context_length_exceeded before the model is called (on /v1/chat/completions, /v1/responses and /v1/messages); a prompt close to the limit is sent to the model, and if the model finds it too long you get the same 400 context_length_exceeded (as an error event mid-stream when stream: true). Either way nothing is billed.

Prompt caching

Models that support prompt caching reuse repeated prompt prefixes automatically — there is nothing to turn on. Cached prompt tokens are billed at the model's cached-input rate where one is published (see Models); a model with no cached rate bills them at its normal input rate. Tokens written to the cache are billed at the input rate, except on models that publish a separate cache-write rate (such as gpt-6.1-sol).

Every response reports the cached share of the prompt:

"usage": {
  "prompt_tokens": 3120,
  "completion_tokens": 42,
  "total_tokens": 3162,
  "prompt_tokens_details": { "cached_tokens": 2944 }
}
  • prompt_tokens is the whole prompt, cached part included.
  • prompt_tokens_details.cached_tokens is always present (0 on a miss or the first request).
  • prompt_tokens_details.cache_write_tokens appears when the model reported tokens written to the cache (GPT-5.6 and later bill these at their own rate).
  • Streaming: the same block is in the final usage chunk when you send stream_options: {"include_usage": true}.
  • /v1/responses reports the same numbers as usage.input_tokens_details.cached_tokens and usage.input_tokens_details.cache_write_tokens.

To improve hit rates:

  • Put stable content first (system prompt, tool definitions, reference documents) and the changing part last.
  • Send a prompt_cache_key (up to 1,024 characters, on /v1/chat/completions and /v1/responses) and reuse it for requests that share a prefix. It is passed on as a cache-routing hint for kimi-k2.5, kimi-k2.6, kimi-k2.7-code, glm-4.7-flash, glm-5.2, gpt-oss-120b, nemotron-3-super, gemma-4-26b-a4b-it, mistral-small-3.1, DeepSeek-V4-Pro and DeepSeek-V4-Flash; other models ignore it.
  • Explicit per-block cache breakpoints (cache_control / prompt_cache_breakpoint on a content part) are not applied today — caching works from the prompt prefix automatically.
{
  "model": "kimi-k2.6",
  "prompt_cache_key": "support-bot:policy-v3",
  "messages": [
    {"role": "system", "content": "<long, stable instructions>"},
    {"role": "user", "content": "Where is my order?"}
  ]
}

A prefix usually needs to be at least ~1,024 tokens before it is cached, and an idle cache expires after a few minutes. Hits are not guaranteed.

Responses API

For clients built on OpenAI's newer Responses API, CallMissed exposes a compatible POST /v1/responses endpoint. It accepts a Responses-shaped body and translates to the same chat engine under the hood — so you can point an OpenAI Responses client at https://api.callmissed.com/v1 without changes.

Endpoint: POST /v1/responses

from openai import OpenAI

client = OpenAI(api_key="cm_your_key", base_url="https://api.callmissed.com/v1")

resp = client.responses.create(
    model="gpt-4.1",
    input="Write a haiku about databases.",
)
print(resp.output_text)
FieldTypeNotes
modelstringAny chat model id. Defaults to kimi-k2.5 when omitted
inputstring or arrayA plain string, or the Responses item array (messages, function_call and function_call_output items). Message content parts: input_text, input_image (image_url), input_file (inline file_data). A function_call_output output may be a string or an array of those parts; its images are passed to the model with the tool result
instructionsstringSystem prompt
max_output_tokensinteger1–1,048,576
temperature / top_pnumber0–2 / 0–1
toolsarrayFlat function tools ({"type": "function", "name", "description", "parameters"}). Any other tool type returns 400 unsupported_tool_type
tool_choice, parallel_tool_callsAs in the Responses API. A tool_choice naming a hosted tool returns 400 unsupported_tool_choice
reasoningobject{"effort": "..."} — mapped like reasoning_effort on chat completions
textobject{"format": {"type": "json_object" | "json_schema", ...}} for structured output
streambooleanEmits Responses-style SSE events (response.output_text.delta, …)
user, provider, prompt_cache_keySame meaning as on /v1/chat/completions
  • The same models, pricing, plan rules, vision and tool-calling support as /v1/chat/completions apply — this is a request/response-shape adapter, not a different model set.
  • Responses are not stored: store is accepted for compatibility and ignored, and previous_response_id returns 400 unsupported_parameter. Send the full conversation in input each turn.
  • An input_image given only as file_id, or an input_file given as file_id or file_url, returns 400 unsupported_content_part naming the part. Send the bytes inline instead.

If you're starting fresh, /v1/chat/completions is the most widely-supported surface; use /v1/responses when porting an existing Responses-API integration.

Errors

All errors return the OpenAI-compatible envelope:

{
  "error": {
    "message": "Invalid API key",
    "type": "invalid_request_error",
    "code": "invalid_api_key"
  }
}
StatuscodeWhen
400unsupported_image_inputImage content sent to a model without vision support
400unsupported_audio_input / unsupported_file_inputinput_audio or file parts sent to a model that doesn't take them
400unsupported_content_partA content part that can't be passed on as sent (for example a file_id reference on /v1/responses)
400voice_agent_only_modelA realtime or managed-voice model id — use POST /v1/voice/sessions
400zdr_unavailableZero data retention is on and the model has no zero-retention route
400context_length_exceededThe prompt is longer than the model's context window
400max_tokens_too_smallmax_tokens below 3 on gpt-5.6-luna, gpt-6-sol, gpt-6-luna or gpt-6.1-sol, or below 16 on gpt-6.1-sol when tools are sent
400max_tokens_too_largemax_tokens above the model's max_output_tokens (for example 65,536 on the gemini-* models, 16,384 on gpt-4o, 32,768 on gpt-4.1, 128,000 on the GPT-5 and GPT-6 models)
401invalid_api_key / api_key_expiredMissing, malformed or revoked key / expired key
402insufficient_creditsBalance exhausted (X-Credits-Balance carries the balance)
402budget_exceededThe key's own budget cap was reached
402end_user_budget_exceededThe user id has used its monthly budget
403permission_deniedThe key lacks the llm permission
403account_inactiveThe account is suspended or closed
403model_not_availableA free-plan key calling a paid model
403model_not_allowedThe key's allowed_models list excludes the model
404model_not_foundUnknown model id
422—Request body failed validation (a field out of range, a missing messages). Body is {"detail": [{loc, msg, type}]}
429quota_exceededMonthly plan call cap reached
429rate_limit_exceededPer-key requests-per-minute exceeded
429too_many_concurrent_requestsToo many requests in flight on this key. Honour Retry-After
503model_under_maintenanceThe model is temporarily unavailable; the message names an alternative
503provider_errorThe model is temporarily unavailable upstream. Safe to retry

Every response carries an X-Request-ID header — quote it when you contact support.