Send a single-turn or multi-turn conversation and receive a complete response. Use any OpenAI SDK — set base_url to https://api.callmissed.com/v1 and api_key to your cm_ key.
from openai import OpenAIclient = OpenAI( api_key="cm_your_key", base_url="https://api.callmissed.com/v1")response = client.chat.completions.create( model="sarvam-105b", messages=[ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "What is the capital of India?"} ])print(response.choices[0].message.content)
Model ID (e.g. sarvam-105b, gpt-5.6-luna). Defaults to sarvam-105b when omitted
messages
array
Required. 1–2,000 {role, content} objects. System prompt goes here as {"role": "system", "content": "..."}. content may be a string or an array of text, image, audio or document parts (see Vision and Audio and documents)
stream
boolean
Enable streaming SSE responses (default false)
stream_options
object
{"include_usage": true} to get token counts in the stream
temperature
number
Sampling temperature, 0–2
max_tokens
integer
Maximum tokens to generate, 1–1,048,576. A model with a documented output limit (the max_output_tokens field in GET /api/v1/models) rejects a larger value with 400 max_tokens_too_large
max_completion_tokens
integer
Same as max_tokens (the newer OpenAI name); used when max_tokens is absent
n
integer
Number of completions, 1–5 (default 1)
top_p
number
Nucleus sampling, 0–1
top_k
integer
Top-K sampling, 1–200
frequency_penalty
number
Penalize repeated tokens, -2–2
presence_penalty
number
Penalize new topics, -2–2
repetition_penalty
number
Reduce repetition, 0–2
seed
integer
Deterministic sampling, 0–2^63-1
stop
string or array
Up to 16 stop sequences
logit_bias
object
Token id → bias between -100 and 100; at most 1,024 entries
"auto", "none", "required", or {"type": "function", "function": {"name": "..."}}
parallel_tool_calls
boolean
Allow parallel function calls
response_format
object
{"type": "json_object"} or {"type": "json_schema", "json_schema": {...}}
structured_outputs
boolean
Enforce strict JSON schema
reasoning_effort
string
"none" / "minimal" / "low" / "medium" / "high" / "xhigh" — each model accepts a different subset and the gateway maps the rest; see the per-model matrix
models
array
Up to 32 fallback model ids, tried in order if model fails with a retryable error — see Model substitution
user
string
Your end user's id. Enforces that user's monthly budget when one is set (caps apply to ids of at most 256 characters)
Up to 1,024 characters. Reuse the same key for requests that share a long prompt prefix to raise the cache hit rate — see Prompt caching
bot_id
string
ID of one of your bots. Its knowledge base is searched with the latest user message and the top matches are added to the prompt — see Knowledge
knowledge_top_k
integer
With bot_id: how many knowledge chunks to add, 1–50 (default 6)
knowledge_min_score
number
With bot_id: minimum similarity score, 0–1 (default 0)
On GPT-5 and GPT-6 family models, temperature, top_p and logit_bias are not supported by the model and are left out of the request; the model runs at its default sampling.
OpenAI Python SDK note — The OpenAI client validates kwargs against its
known parameters, so a CallMissed-specific field such as reasoning_effort
raises TypeError: Completions.create() got an unexpected keyword argument.
Pass it via extra_body instead:
If model fails with a retryable error (an upstream outage or rate limit), the
next id in models is tried. Each fallback must pass the same checks as
model — your plan, the key's allowed_models, maintenance status, vision
support and context window — and ids that fail them are skipped. The response's
model field names the model that actually answered, and you are billed at
that model's rate. Requests that send tools, tool_choice,
response_format or structured_outputs never fall back, because a different
model could change the result shape.
Multimodal content (text + image parts) is accepted on any model whose
supports_vision flag is true in GET /v1/models. Models without vision
support reject image content with 400 unsupported_image_inputbefore the
upstream call, so you're not charged.
from openai import OpenAIclient = OpenAI(api_key="cm_your_key", base_url="https://api.callmissed.com/v1")resp = client.chat.completions.create( model="gpt-5.6-sol", # supports_vision: true messages=[{ "role": "user", "content": [ {"type": "text", "text": "What is in this image?"}, {"type": "image_url", "image_url": {"url": "https://example.com/cat.png"}}, ], }],)
GET /v1/models is authoritative. Read supports_vision there rather than
hard-coding this list.
An image_url may be an https:// URL or a base64 data URL. The Gemini models
and gemma-4-31b only take image bytes, so CallMissed downloads a URL image
for them: it must be publicly reachable without redirects, return an image
content type, and be at most 20 MB (10 MB for gemma-4-31b). Gemini accepts
PNG, JPEG, WEBP, HEIC and HEIF. If a download fails the request returns 400;
send base64 to avoid the extra fetch.
The Gemini models (gemini-3.8-flash, gemini-3.7-flash, gemini-3.6-flash,
gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.1-pro-preview,
gemini-3.1-flash-lite) also take audio and PDF parts:
Audio format may be wav, mp3, aac, flac, ogg, opus, m4a, aiff
or webm. A file part must carry the PDF inline in file_data; file_id
references are not supported. Any other model rejects these parts with
400 unsupported_audio_input or 400 unsupported_file_input before the
upstream call, so you're not charged. Content that can't be passed on as sent
is rejected with a 400 that names the part, rather than being dropped.
Every model in the catalog advertises a context_window (token count for the
combined prompt + completion). The GET /v1/models response exposes it under
two keys for cross-client compatibility:
context_window (OpenAI/CallMissed canonical name)
context_length (OpenAI SDK convention — same value)
from openai import OpenAIclient = OpenAI(api_key="cm_your_key", base_url="https://api.callmissed.com/v1")for m in client.models.list(): extra = m.model_extra or {} print(m.id, extra.get("context_window"), extra.get("supports_vision"))
Images, audio and files count toward the window by what the model reads, not
by their encoded size, so a large base64 image does not by itself cause
context_length_exceeded. A prompt that is clearly longer than the window is
rejected with 400 context_length_exceeded before the model is called (on
/v1/chat/completions, /v1/responses and /v1/messages); a prompt close to
the limit is sent to the model, and if the model finds it too long you get the
same 400 context_length_exceeded (as an error event mid-stream when
stream: true). Either way nothing is billed.
Models that support prompt caching reuse repeated prompt prefixes
automatically — there is nothing to turn on. Cached prompt tokens are billed at
the model's cached-input rate where one is published (see Models);
a model with no cached rate bills them at its normal input rate. Tokens written
to the cache are billed at the input rate, except on models that publish a
separate cache-write rate (such as gpt-6.1-sol).
Every response reports the cached share of the prompt:
prompt_tokens is the whole prompt, cached part included.
prompt_tokens_details.cached_tokens is always present (0 on a miss or the first request).
prompt_tokens_details.cache_write_tokens appears when the model reported tokens written to the cache (GPT-5.6 and later bill these at their own rate).
Streaming: the same block is in the final usage chunk when you send stream_options: {"include_usage": true}.
/v1/responses reports the same numbers as usage.input_tokens_details.cached_tokens and usage.input_tokens_details.cache_write_tokens.
To improve hit rates:
Put stable content first (system prompt, tool definitions, reference documents) and the changing part last.
Send a prompt_cache_key (up to 1,024 characters, on /v1/chat/completions and /v1/responses) and reuse it for requests that share a prefix. It is passed on as a cache-routing hint for kimi-k2.5, kimi-k2.6, kimi-k2.7-code, glm-4.7-flash, glm-5.2, gpt-oss-120b, nemotron-3-super, gemma-4-26b-a4b-it, mistral-small-3.1, DeepSeek-V4-Pro and DeepSeek-V4-Flash; other models ignore it.
Explicit per-block cache breakpoints (cache_control / prompt_cache_breakpoint on a content part) are not applied today — caching works from the prompt prefix automatically.
For clients built on OpenAI's newer Responses API, CallMissed exposes a compatible POST /v1/responses endpoint. It accepts a Responses-shaped body and translates to the same chat engine under the hood — so you can point an OpenAI Responses client at https://api.callmissed.com/v1 without changes.
Endpoint:POST /v1/responses
from openai import OpenAIclient = OpenAI(api_key="cm_your_key", base_url="https://api.callmissed.com/v1")resp = client.responses.create( model="gpt-4.1", input="Write a haiku about databases.",)print(resp.output_text)
Field
Type
Notes
model
string
Any chat model id. Defaults to kimi-k2.5 when omitted
input
string or array
A plain string, or the Responses item array (messages, function_call and function_call_output items). Message content parts: input_text, input_image (image_url), input_file (inline file_data). A function_call_outputoutput may be a string or an array of those parts; its images are passed to the model with the tool result
instructions
string
System prompt
max_output_tokens
integer
1–1,048,576
temperature / top_p
number
0–2 / 0–1
tools
array
Flat function tools ({"type": "function", "name", "description", "parameters"}). Any other tool type returns 400 unsupported_tool_type
tool_choice, parallel_tool_calls
As in the Responses API. A tool_choice naming a hosted tool returns 400 unsupported_tool_choice
reasoning
object
{"effort": "..."} — mapped like reasoning_effort on chat completions
text
object
{"format": {"type": "json_object" | "json_schema", ...}} for structured output
The same models, pricing, plan rules, vision and tool-calling support as /v1/chat/completions apply — this is a request/response-shape adapter, not a different model set.
Responses are not stored: store is accepted for compatibility and ignored, and previous_response_id returns 400 unsupported_parameter. Send the full conversation in input each turn.
An input_image given only as file_id, or an input_file given as file_id or file_url, returns 400 unsupported_content_part naming the part. Send the bytes inline instead.
If you're starting fresh, /v1/chat/completions is the most widely-supported surface; use /v1/responses when porting an existing Responses-API integration.
Image content sent to a model without vision support
400
unsupported_audio_input / unsupported_file_input
input_audio or file parts sent to a model that doesn't take them
400
unsupported_content_part
A content part that can't be passed on as sent (for example a file_id reference on /v1/responses)
400
voice_agent_only_model
A realtime or managed-voice model id — use POST /v1/voice/sessions
400
zdr_unavailable
Zero data retention is on and the model has no zero-retention route
400
context_length_exceeded
The prompt is longer than the model's context window
400
max_tokens_too_small
max_tokens below 3 on gpt-5.6-luna, gpt-6-sol, gpt-6-luna or gpt-6.1-sol, or below 16 on gpt-6.1-sol when tools are sent
400
max_tokens_too_large
max_tokens above the model's max_output_tokens (for example 65,536 on the gemini-* models, 16,384 on gpt-4o, 32,768 on gpt-4.1, 128,000 on the GPT-5 and GPT-6 models)
401
invalid_api_key / api_key_expired
Missing, malformed or revoked key / expired key
402
insufficient_credits
Balance exhausted (X-Credits-Balance carries the balance)
402
budget_exceeded
The key's own budget cap was reached
402
end_user_budget_exceeded
The user id has used its monthly budget
403
permission_denied
The key lacks the llm permission
403
account_inactive
The account is suspended or closed
403
model_not_available
A free-plan key calling a paid model
403
model_not_allowed
The key's allowed_models list excludes the model
404
model_not_found
Unknown model id
422
—
Request body failed validation (a field out of range, a missing messages). Body is {"detail": [{loc, msg, type}]}
429
quota_exceeded
Monthly plan call cap reached
429
rate_limit_exceeded
Per-key requests-per-minute exceeded
429
too_many_concurrent_requests
Too many requests in flight on this key. Honour Retry-After
503
model_under_maintenance
The model is temporarily unavailable; the message names an alternative
503
provider_error
The model is temporarily unavailable upstream. Safe to retry
Every response carries an X-Request-ID header — quote it when you contact support.