Skip to main content

Kimi K2.5 Fast (Maintenance)

High-throughput Kimi K2.5 inference tier — currently under maintenance. Use kimi-k2.5 in the meantime.

Under maintenance. kimi-k2.5-fast is temporarily unavailable. Requests return HTTP 503 with code: "model_under_maintenance". Use kimi-k2.5 for production traffic; both ride on the same Kimi K2.5 model from Moonshot AI.

Overview

The kimi-k2.5-fast tier targets ultra-low-latency voice-agent workloads on a dedicated high-throughput serving tier. While it's under maintenance, route the same workload through kimi-k2.5 — the model and tokeniser are identical, only the inference latency differs.

Kimi K2.5 Fast

FieldValue
Model IDkimi-k2.5-fast
StatusUnder maintenance — returns 503
Recommended fallbackkimi-k2.5
ArchitectureMoE (Mixture of Experts)
Context window256,000 tokens
Supports streamingYes
Supports toolsYes

Kimi K2.5 (by Moonshot AI) is a 1T-parameter MoE model with 32B active parameters. It excels at reasoning, coding, and multilingual tasks.

Usage

While kimi-k2.5-fast is in maintenance, point your code at kimi-k2.5:

from openai import OpenAI

client = OpenAI(
    base_url="https://api.callmissed.com/v1",
    api_key="cm_your_api_key",
)

response = client.chat.completions.create(
    model="kimi-k2.5",  # kimi-k2.5-fast is under maintenance
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain quantum computing briefly."},
    ],
    stream=True,
)

for chunk in response:
    print(chunk.choices[0].delta.content or "", end="", flush=True)

Pricing

DirectionCost per 1M tokens
Input$0.8438
Output$4.219

1 credit = ₹1 ≈ US$0.0104 (US$1 = ₹96). A typical voice-agent turn — 500 input + 200 output tokens — costs $0.001266, or 0.1215 credits. These rates apply once kimi-k2.5-fast leaves maintenance. See Credits & Rate Limits.