Skip to main content

Embeddings

Turn text into vectors with the OpenAI-compatible embeddings endpoint — batching, dimensions, base64 output, and per-token pricing.

Overview

POST /v1/embeddings converts text into a dense float vector you can store in your own vector database and search with cosine similarity. It is the primitive behind retrieval, semantic search, clustering, deduplication and classification.

The request and response are OpenAI-compatible, so the official OpenAI SDKs work unchanged once you point them at https://api.callmissed.com/v1 with a cm_ key.

If you want retrieval without running your own vector store, use Knowledge & RAG instead — it ingests, chunks, embeds and searches for you.

Authentication

Authorization: Bearer cm_your_api_key

This endpoint is gated by the key's service permission, not by a resource scope. The key needs llm (or *). There is no separate embedding permission — a key that can call /v1/chat/completions can call /v1/embeddings.

A key without it returns 403:

{
  "error": {
    "message": "This API key does not have permission for embeddings (requires LLM permission). Update key permissions in your dashboard.",
    "type": "invalid_request_error",
    "code": "permission_denied"
  }
}

Models

Both embedding models are free-plan callable — they are metered per input token, so your credit balance is the only governor.

ModelDimensionsMax inputPrice (per 1M input tokens)
text-embedding-3-small15368,192 tokens$0.02083
text-embedding-3-large30728,192 tokens$0.1354

Start with text-embedding-3-small: it is the better price/performance choice for large corpora. Move to -large only when you have measured that retrieval quality is the bottleneck.

Both appear in GET /v1/models with "owned_by": "openai".

Quickstart

curl https://api.callmissed.com/v1/embeddings \
  -H "Authorization: Bearer cm_your_api_key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "text-embedding-3-small",
    "input": "Where is my order?"
  }'
{
  "object": "list",
  "data": [
    {
      "object": "embedding",
      "index": 0,
      "embedding": [0.0023064255, -0.009327292, 0.015797347]
    }
  ],
  "model": "text-embedding-3-small",
  "usage": {
    "prompt_tokens": 5,
    "total_tokens": 5
  }
}
from openai import OpenAI

client = OpenAI(
    api_key="cm_your_api_key",
    base_url="https://api.callmissed.com/v1",
)

resp = client.embeddings.create(
    model="text-embedding-3-small",
    input=["Where is my order?", "How do I return this?"],
)
vectors = [row.embedding for row in resp.data]

Request

FieldTypeRequiredDefaultConstraints
modelstringYes—text-embedding-3-small or text-embedding-3-large
inputstring or string[]Yes—Up to 128 items per request; each item non-empty, at most 100,000 characters and within the model's 8,192-token limit. Pre-tokenised integer arrays are not accepted
encoding_formatstringNofloatfloat or base64
dimensionsintegerNomodel native1 <= dimensions <= 3072 (large) or 1536 (small)
userstringNo—At most 256 characters. Your end user's id; enforces that user's monthly budget when one is set
providerobjectNo—{"zdr": true} requires a zero-data-retention route; otherwise 400 zdr_unavailable

Batching

Send an array to embed up to 128 strings in one round trip. The index on each row matches the position in your input array, so you can zip the results back onto your records without re-ordering.

{
  "model": "text-embedding-3-small",
  "input": ["first chunk", "second chunk", "third chunk"]
}

A batch of 129 or more returns 422:

{ "error": { "message": "`input` array too long: 200 items (maximum 128). Split the batch across multiple requests.", "type": "invalid_request_error", "code": "invalid_request_error" } }

Shortening vectors with dimensions

Both models support Matryoshka-style truncation. Passing dimensions returns a shorter, renormalised vector — smaller index, faster search, slightly lower recall.

{ "model": "text-embedding-3-large", "input": "hello", "dimensions": 256 }

dimensions must be between 1 and the model's native size. Anything else returns 422.

Vectors of different lengths are not comparable. Pick one model and one dimensions value per index and keep it fixed — re-embed the whole corpus if you change either.

encoding_format: "base64"

base64 returns each vector as a base64 string of little-endian float32 values instead of a JSON array. It is roughly a third of the payload size, which matters when you are embedding thousands of chunks.

import base64, struct

raw = base64.b64decode(resp.data[0].embedding)
vector = list(struct.unpack(f"<{len(raw) // 4}f", raw))

Response

FieldTypeNotes
objectstringAlways list
data[].objectstringAlways embedding
data[].indexintegerPosition in your input array
data[].embeddingnumber[] or stringFloat array, or a base64 string when encoding_format is base64
modelstringThe model that served the request
usage.prompt_tokensintegerInput tokens billed
usage.total_tokensintegerSame as prompt_tokens — embeddings have no output tokens

Billing

Embeddings are metered on input tokens only. Credits are deducted as tokens / 1,000,000 x rate in credits, where 1 credit = ₹1 ≈ US$0.0104 (US$1 = ₹96).

  • text-embedding-3-small — 2 credits per 1M input tokens
  • text-embedding-3-large — 13 credits per 1M input tokens

A request that fails with a 4xx or 5xx is recorded in your usage log but not charged. Track spend with GET /v1/usage/summary.

Errors

StatuscodeWhen
400invalid_request_errorBody is not a JSON object
400context_length_exceededAn input item exceeds the model's 8,192-token limit
401invalid_api_key / api_key_expiredMissing, malformed, or expired key
402insufficient_creditsBalance is exhausted. The X-Credits-Balance header carries the current balance
402budget_exceededThe key's own budget cap was hit
403permission_deniedKey lacks the llm permission
403model_not_allowedThe key's allowed_models list excludes this model
404model_not_foundUnknown embedding model id
413invalid_request_errorAn input item is over 100,000 characters
422invalid_request_errorBatch too long, empty input, bad dimensions, unsupported encoding_format, or a token array instead of a string
429rate_limit_exceededPer-key request rate exceeded. Retry with backoff
429quota_exceededPlan or monthly budget cap reached. Honour Retry-After
502upstream_errorEmbedding generation failed. Safe to retry
503service_unavailableTemporary capacity problem. Retry with backoff

Every error uses the standard envelope:

{ "error": { "message": "…", "type": "invalid_request_error", "code": "model_not_found", "request_id": "req_…" } }

Building a search index

  1. Chunk your documents to roughly 200–500 tokens with a little overlap.
  2. Embed chunks in batches of 128 with text-embedding-3-small.
  3. Store {id, text, vector, metadata} in your vector database.
  4. At query time, embed the query with the same model and dimensions, then retrieve by cosine similarity.
  5. Pass the top chunks to POST /v1/chat/completions as context.
query = client.embeddings.create(
    model="text-embedding-3-small",
    input="refund policy",
).data[0].embedding
# hits = your_vector_db.search(query, top_k=5)