Anthropic-Compatible API
Use the Anthropic SDK with CallMissed — just change the base URL. Full Messages API compatibility.
Overview
CallMissed provides an Anthropic Messages API-compatible endpoint alongside the OpenAI-compatible API. If you're already using the Anthropic SDK, you can switch to CallMissed by changing only the base_url.
Endpoints:
POST /v1/messages— chat completions (streaming + non-streaming)POST /v1/messages/count_tokens— token count estimation (real BPE, not char-based); also at/anthropic/v1/messages/count_tokensGET /anthropic/v1/models— list models in Anthropic shape with capability metadataGET /anthropic/v1/models/{model_id}— single model detailPOST /anthropic/v1/messages— alternate path for the chat endpoint
Authentication: Use either header style:
x-api-key: cm_your_key(Anthropic SDK default)Authorization: Bearer cm_your_key(OpenAI style)
Basic Usage
import anthropic
client = anthropic.Anthropic(
api_key="cm_your_key",
base_url="https://api.callmissed.com"
)
message = client.messages.create(
model="gpt-5.6-sol",
max_tokens=1024,
system="You are a helpful assistant.",
messages=[
{"role": "user", "content": "What is the capital of India?"}
]
)
print(message.content[0].text)Response:
{
"id": "msg-abc123def456",
"type": "message",
"role": "assistant",
"content": [
{"type": "text", "text": "The capital of India is New Delhi."}
],
"model": "gpt-5.6-sol",
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 25,
"output_tokens": 12,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 0
}
}Streaming
Set stream: true to receive Server-Sent Events with the full Anthropic streaming lifecycle:
import anthropic
client = anthropic.Anthropic(
api_key="cm_your_key",
base_url="https://api.callmissed.com"
)
with client.messages.stream(
model="gpt-5.6-sol",
max_tokens=1024,
messages=[{"role": "user", "content": "Tell me a short story."}]
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)SSE event lifecycle:
event: message_start → message metadata + input token count
event: content_block_start → new content block begins
event: content_block_delta → text chunks (repeats)
event: content_block_stop → content block complete
event: message_delta → stop_reason + output token count
event: message_stop → stream completeParameters
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID (e.g. gpt-5.6-sol, sarvam-105b, kimi-k2.6) |
max_tokens | integer | Yes | Maximum tokens to generate, 1–1,048,576 |
messages | array | Yes | 1–2,000 {role, content} objects. content is a string or an array of text, image, tool_use and tool_result blocks |
system | string or array | No | System prompt (top-level, not in messages); a string or a list of text blocks |
stream | boolean | No | Enable streaming (default: false) |
temperature | number | No | Sampling temperature, 0–1 |
top_p | number | No | Nucleus sampling, 0–1 |
top_k | integer | No | Top-K sampling, 1–500 |
stop_sequences | array | No | Up to 16 stop sequences |
tools | array | No | Up to 128 {name, description, input_schema} tools |
tool_choice | object | No | {"type": "auto"}, {"type": "any"}, {"type": "none"} or {"type": "tool", "name": "..."} |
metadata | object | No | Up to 16 string, number or boolean values. user_id is your end user's id and enforces that user's monthly budget. trace_id and session_id are recorded on the usage row so you can filter usage logs by them |
provider | object | No | CallMissed extension: {"zdr": true} requires a zero-data-retention route |
Note: Unlike the OpenAI API,
max_tokensis required andsystemis a top-level parameter (not a message withrole: "system").
Model Selection
Send any model ID from the Models catalog — not just Anthropic-shaped names. The model field takes the same values as /v1/chat/completions.
{ "model": "gpt-5.6-sol", "max_tokens": 1024, "messages": [...] }Token Counting
Estimate input token count before sending a request. The endpoint uses a BPE
tokenizer (tiktoken cl100k_base) — close to Claude's real tokenizer on
typical English prompts, and noticeably more accurate than char-length
heuristics.
curl -X POST https://api.callmissed.com/v1/messages/count_tokens \
-H "x-api-key: cm_your_key" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"messages": [{"role": "user", "content": "Hello, how are you?"}],
"system": "You are a helpful assistant."
}'Response:
{"input_tokens": 15}Image content blocks contribute a fixed estimate (~258 tokens per image) rather than a fetch-and-resize pass. Tool definitions are counted against the total by serializing each to JSON and tokenizing the schema.
Listing Models
List all available models via the Anthropic-shape endpoint:
curl https://api.callmissed.com/anthropic/v1/models \
-H "x-api-key: cm_your_key"Response:
{
"data": [
{
"type": "model",
"id": "gpt-5.6-sol",
"display_name": "GPT-5.6 Sol",
"created_at": "2023-11-14T22:13:20+00:00",
"description": "Frontier model for complex professional work. Multimodal, reasoning + tools.",
"category": "llm",
"context_window": 1050000,
"context_length": 1050000,
"pricing": {"input": 5.208, "output": 31.25, "unit": "per_million_tokens", "currency": "USD"},
"supports_streaming": true,
"supports_tools": true,
"supports_reasoning": true,
"supports_vision": true
}
],
"has_more": false,
"first_id": "...",
"last_id": "..."
}Fetch a single model at GET /anthropic/v1/models/{model_id}.
Vision (Image Input)
Send images on any model whose supports_vision flag is true in the model
listing. That is the authoritative source; see the
vision list for the current set.
Models without vision reject image content with 400 invalid_request_error
before the upstream call — you are not charged.
curl -X POST https://api.callmissed.com/v1/messages \
-H "x-api-key: cm_your_key" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"max_tokens": 1024,
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": "<base64>"}}
]
}]
}'Error Format
Errors return the Anthropic format (different from the OpenAI endpoints):
{
"type": "error",
"error": {
"type": "authentication_error",
"message": "Invalid API key"
}
}| Error type | HTTP Status | When |
|---|---|---|
invalid_request_error | 400 / 422 | Bad request, image sent to a text-only model, or a voice-only model id |
authentication_error | 401 | Bad, missing, revoked or expired API key |
billing_error | 402 | Insufficient credits, key budget or end-user budget exhausted |
permission_error | 403 | Account inactive, key lacks the llm permission, or a free-plan key calling a paid model |
not_found_error | 404 | Model not found |
request_too_large | 413 | Request body too large |
rate_limit_error | 429 | Plan limit or API key rate limit exceeded |
api_error | 500 / 502 / 503 | Upstream model failure, or the model is under maintenance (the message names an alternative) |
overloaded_error | 503 / 529 | Model temporarily unavailable. Retry with backoff |
timeout_error | 504 | Upstream model timed out |
Rate limit headers are returned in Anthropic format:
anthropic-ratelimit-requests-limit: 60
anthropic-ratelimit-requests-remaining: 45
anthropic-ratelimit-requests-reset: 2026-05-01T00:00:00+00:00Prompt caching
Models that support prompt caching reuse repeated prompt prefixes automatically. Usage uses Anthropic's field names, with the same meaning:
input_tokens— prompt tokens that were not read from or written to the cachecache_read_input_tokens— prompt tokens served from the cache (billed at the model's cached-input rate)cache_creation_input_tokens— prompt tokens written to the cache
Total prompt = input_tokens + cache_read_input_tokens + cache_creation_input_tokens.
When streaming, the final counts arrive in the message_delta event.
cache_control blocks (on system, message content or tools) are accepted
so Anthropic SDK code runs unchanged, but explicit breakpoints and ttl are not
applied today — caching works from the prompt prefix automatically.
Differences from Anthropic
This endpoint is designed to work with the Anthropic SDK out of the box. Key differences from the official Anthropic API:
anthropic-versionheader is accepted but not required- Model routing — requests can target any model in the CallMissed catalogue, not just Anthropic-shaped names.
- Token counting uses a BPE tokenizer approximation (tiktoken
cl100k_base). Expect ~5-10% variance from Anthropic's native counts on English prompts; larger on CJK and heavy-punctuation text. - Tools are supported —
toolsandtool_choicework as documented, andtool_use/tool_resultcontent blocks are preserved. - Vision is supported on models whose
supports_visionflag istrue. Image content sent to text-only models is rejected with a400 invalid_request_errorbefore the upstream call, so your credits are safe. - Message Batches API (
/v1/messages/batches) is not implemented — use the regular/v1/messagesendpoint. - Billing uses CallMissed credits, not Anthropic billing.