Skip to main content

Anthropic-Compatible API

Use the Anthropic SDK with CallMissed — just change the base URL. Full Messages API compatibility.

Overview

CallMissed provides an Anthropic Messages API-compatible endpoint alongside the OpenAI-compatible API. If you're already using the Anthropic SDK, you can switch to CallMissed by changing only the base_url.

Endpoints:

  • POST /v1/messages — chat completions (streaming + non-streaming)
  • POST /v1/messages/count_tokens — token count estimation (real BPE, not char-based); also at /anthropic/v1/messages/count_tokens
  • GET /anthropic/v1/models — list models in Anthropic shape with capability metadata
  • GET /anthropic/v1/models/{model_id} — single model detail
  • POST /anthropic/v1/messages — alternate path for the chat endpoint

Authentication: Use either header style:

  • x-api-key: cm_your_key (Anthropic SDK default)
  • Authorization: Bearer cm_your_key (OpenAI style)

Basic Usage

import anthropic

client = anthropic.Anthropic(
    api_key="cm_your_key",
    base_url="https://api.callmissed.com"
)

message = client.messages.create(
    model="gpt-5.6-sol",
    max_tokens=1024,
    system="You are a helpful assistant.",
    messages=[
        {"role": "user", "content": "What is the capital of India?"}
    ]
)

print(message.content[0].text)

Response:

{
  "id": "msg-abc123def456",
  "type": "message",
  "role": "assistant",
  "content": [
    {"type": "text", "text": "The capital of India is New Delhi."}
  ],
  "model": "gpt-5.6-sol",
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 25,
    "output_tokens": 12,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 0
  }
}

Streaming

Set stream: true to receive Server-Sent Events with the full Anthropic streaming lifecycle:

import anthropic

client = anthropic.Anthropic(
    api_key="cm_your_key",
    base_url="https://api.callmissed.com"
)

with client.messages.stream(
    model="gpt-5.6-sol",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Tell me a short story."}]
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

SSE event lifecycle:

event: message_start        → message metadata + input token count
event: content_block_start  → new content block begins
event: content_block_delta  → text chunks (repeats)
event: content_block_stop   → content block complete
event: message_delta        → stop_reason + output token count
event: message_stop          → stream complete

Parameters

ParameterTypeRequiredDescription
modelstringYesModel ID (e.g. gpt-5.6-sol, sarvam-105b, kimi-k2.6)
max_tokensintegerYesMaximum tokens to generate, 1–1,048,576
messagesarrayYes1–2,000 {role, content} objects. content is a string or an array of text, image, tool_use and tool_result blocks
systemstring or arrayNoSystem prompt (top-level, not in messages); a string or a list of text blocks
streambooleanNoEnable streaming (default: false)
temperaturenumberNoSampling temperature, 0–1
top_pnumberNoNucleus sampling, 0–1
top_kintegerNoTop-K sampling, 1–500
stop_sequencesarrayNoUp to 16 stop sequences
toolsarrayNoUp to 128 {name, description, input_schema} tools
tool_choiceobjectNo{"type": "auto"}, {"type": "any"}, {"type": "none"} or {"type": "tool", "name": "..."}
metadataobjectNoUp to 16 string, number or boolean values. user_id is your end user's id and enforces that user's monthly budget. trace_id and session_id are recorded on the usage row so you can filter usage logs by them
providerobjectNoCallMissed extension: {"zdr": true} requires a zero-data-retention route

Note: Unlike the OpenAI API, max_tokens is required and system is a top-level parameter (not a message with role: "system").

Model Selection

Send any model ID from the Models catalog — not just Anthropic-shaped names. The model field takes the same values as /v1/chat/completions.

{ "model": "gpt-5.6-sol", "max_tokens": 1024, "messages": [...] }

Token Counting

Estimate input token count before sending a request. The endpoint uses a BPE tokenizer (tiktoken cl100k_base) — close to Claude's real tokenizer on typical English prompts, and noticeably more accurate than char-length heuristics.

curl -X POST https://api.callmissed.com/v1/messages/count_tokens \
  -H "x-api-key: cm_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "messages": [{"role": "user", "content": "Hello, how are you?"}],
    "system": "You are a helpful assistant."
  }'

Response:

{"input_tokens": 15}

Image content blocks contribute a fixed estimate (~258 tokens per image) rather than a fetch-and-resize pass. Tool definitions are counted against the total by serializing each to JSON and tokenizing the schema.

Listing Models

List all available models via the Anthropic-shape endpoint:

curl https://api.callmissed.com/anthropic/v1/models \
  -H "x-api-key: cm_your_key"

Response:

{
  "data": [
    {
      "type": "model",
      "id": "gpt-5.6-sol",
      "display_name": "GPT-5.6 Sol",
      "created_at": "2023-11-14T22:13:20+00:00",
      "description": "Frontier model for complex professional work. Multimodal, reasoning + tools.",
      "category": "llm",
      "context_window": 1050000,
      "context_length": 1050000,
      "pricing": {"input": 5.208, "output": 31.25, "unit": "per_million_tokens", "currency": "USD"},
      "supports_streaming": true,
      "supports_tools": true,
      "supports_reasoning": true,
      "supports_vision": true
    }
  ],
  "has_more": false,
  "first_id": "...",
  "last_id": "..."
}

Fetch a single model at GET /anthropic/v1/models/{model_id}.

Vision (Image Input)

Send images on any model whose supports_vision flag is true in the model listing. That is the authoritative source; see the vision list for the current set. Models without vision reject image content with 400 invalid_request_error before the upstream call — you are not charged.

curl -X POST https://api.callmissed.com/v1/messages \
  -H "x-api-key: cm_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "max_tokens": 1024,
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "What is in this image?"},
        {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": "<base64>"}}
      ]
    }]
  }'

Error Format

Errors return the Anthropic format (different from the OpenAI endpoints):

{
  "type": "error",
  "error": {
    "type": "authentication_error",
    "message": "Invalid API key"
  }
}
Error typeHTTP StatusWhen
invalid_request_error400 / 422Bad request, image sent to a text-only model, or a voice-only model id
authentication_error401Bad, missing, revoked or expired API key
billing_error402Insufficient credits, key budget or end-user budget exhausted
permission_error403Account inactive, key lacks the llm permission, or a free-plan key calling a paid model
not_found_error404Model not found
request_too_large413Request body too large
rate_limit_error429Plan limit or API key rate limit exceeded
api_error500 / 502 / 503Upstream model failure, or the model is under maintenance (the message names an alternative)
overloaded_error503 / 529Model temporarily unavailable. Retry with backoff
timeout_error504Upstream model timed out

Rate limit headers are returned in Anthropic format:

anthropic-ratelimit-requests-limit: 60
anthropic-ratelimit-requests-remaining: 45
anthropic-ratelimit-requests-reset: 2026-05-01T00:00:00+00:00

Prompt caching

Models that support prompt caching reuse repeated prompt prefixes automatically. Usage uses Anthropic's field names, with the same meaning:

  • input_tokens — prompt tokens that were not read from or written to the cache
  • cache_read_input_tokens — prompt tokens served from the cache (billed at the model's cached-input rate)
  • cache_creation_input_tokens — prompt tokens written to the cache

Total prompt = input_tokens + cache_read_input_tokens + cache_creation_input_tokens. When streaming, the final counts arrive in the message_delta event.

cache_control blocks (on system, message content or tools) are accepted so Anthropic SDK code runs unchanged, but explicit breakpoints and ttl are not applied today — caching works from the prompt prefix automatically.

Differences from Anthropic

This endpoint is designed to work with the Anthropic SDK out of the box. Key differences from the official Anthropic API:

  • anthropic-version header is accepted but not required
  • Model routing — requests can target any model in the CallMissed catalogue, not just Anthropic-shaped names.
  • Token counting uses a BPE tokenizer approximation (tiktoken cl100k_base). Expect ~5-10% variance from Anthropic's native counts on English prompts; larger on CJK and heavy-punctuation text.
  • Tools are supported — tools and tool_choice work as documented, and tool_use/tool_result content blocks are preserved.
  • Vision is supported on models whose supports_vision flag is true. Image content sent to text-only models is rejected with a 400 invalid_request_error before the upstream call, so your credits are safe.
  • Message Batches API (/v1/messages/batches) is not implemented — use the regular /v1/messages endpoint.
  • Billing uses CallMissed credits, not Anthropic billing.