Audio Translation
Translate audio in any supported language to English text. OpenAI-compatible endpoint.
Overview
Translates speech in any of 23 supported languages to English text. OpenAI-compatible /v1/audio/translations.
Unlike Speech to Text (which transcribes in the original language), this endpoint always outputs English.
Endpoint: POST /v1/audio/translations
Supported input languages (23, with saaras:v3 / saaras:v4): Hindi, Bengali, Tamil, Telugu, Kannada, Malayalam, Marathi, Gujarati, Punjabi, Odia, Assamese, Urdu, Nepali, Konkani, Kashmiri, Sindhi, Sanskrit, Santali, Manipuri, Bodo, Maithili, Dogri, English. The source language is always auto-detected — this endpoint takes no language field.
Only three models translate. saaras:v3, saaras:v4 and whisper-large-v3-turbo return English. Any other STT model accepted here returns a transcript in the original language, billed at that model's rate. Stick to those three for translation.
Basic Usage
from openai import OpenAI
client = OpenAI(
api_key="cm_your_key",
base_url="https://api.callmissed.com/v1"
)
# Translate Hindi audio to English text
with open("hindi_audio.wav", "rb") as f:
translation = client.audio.translations.create(
model="saaras:v3",
file=f,
)
print(translation.text)Response:
{"text": "Hello, how are you? I wanted to discuss the project."}Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
file | file | Yes | Audio file (WAV, MP3, AAC, OGG, FLAC, WebM, M4A), up to 25 MB. An empty file returns 400. The same per-model size and duration limits as transcription apply |
model | string | No | Default saaras:v3. saaras:v4 and whisper-large-v3-turbo also translate to English — see the warning above |
response_format | string | No | json (default), text, or verbose_json. Any other value is answered as json |
temperature | float | No | 0–2. Accepted for OpenAI SDK compatibility; currently not forwarded to the model |
prompt | string | No | Accepted for OpenAI SDK compatibility; currently not forwarded to the model |
Errors match Speech to Text: 400 empty file, 402 insufficient credits, 403 missing stt permission or model not on your plan, 404 unknown model, 422 invalid form field, 429 plan limit, 502 model failure.
Response Formats
json (default)
{"text": "Hello, how are you?"}text
Returns plain text with no JSON wrapping.
verbose_json
duration is the billed audio length in seconds. segments and words are always empty.
{
"task": "translate",
"language": "en",
"duration": 4.52,
"text": "Hello, how are you?",
"segments": [],
"words": []
}Tip: For transcription in the original language (not translated), use Speech to Text instead. For output modes like transliteration or code-mixing, use the
modeparameter on the transcription endpoint.