Skip to main content

Audio Translation

Translate audio in any supported language to English text. OpenAI-compatible endpoint.

Overview

Translates speech in any of 23 supported languages to English text. OpenAI-compatible /v1/audio/translations.

Unlike Speech to Text (which transcribes in the original language), this endpoint always outputs English.

Endpoint: POST /v1/audio/translations

Supported input languages (23, with saaras:v3 / saaras:v4): Hindi, Bengali, Tamil, Telugu, Kannada, Malayalam, Marathi, Gujarati, Punjabi, Odia, Assamese, Urdu, Nepali, Konkani, Kashmiri, Sindhi, Sanskrit, Santali, Manipuri, Bodo, Maithili, Dogri, English. The source language is always auto-detected — this endpoint takes no language field.

Only three models translate. saaras:v3, saaras:v4 and whisper-large-v3-turbo return English. Any other STT model accepted here returns a transcript in the original language, billed at that model's rate. Stick to those three for translation.

Basic Usage

from openai import OpenAI

client = OpenAI(
    api_key="cm_your_key",
    base_url="https://api.callmissed.com/v1"
)

# Translate Hindi audio to English text
with open("hindi_audio.wav", "rb") as f:
    translation = client.audio.translations.create(
        model="saaras:v3",
        file=f,
    )

print(translation.text)

Response:

{"text": "Hello, how are you? I wanted to discuss the project."}

Parameters

ParameterTypeRequiredDescription
filefileYesAudio file (WAV, MP3, AAC, OGG, FLAC, WebM, M4A), up to 25 MB. An empty file returns 400. The same per-model size and duration limits as transcription apply
modelstringNoDefault saaras:v3. saaras:v4 and whisper-large-v3-turbo also translate to English — see the warning above
response_formatstringNojson (default), text, or verbose_json. Any other value is answered as json
temperaturefloatNo0–2. Accepted for OpenAI SDK compatibility; currently not forwarded to the model
promptstringNoAccepted for OpenAI SDK compatibility; currently not forwarded to the model

Errors match Speech to Text: 400 empty file, 402 insufficient credits, 403 missing stt permission or model not on your plan, 404 unknown model, 422 invalid form field, 429 plan limit, 502 model failure.

Response Formats

json (default)

{"text": "Hello, how are you?"}

text

Returns plain text with no JSON wrapping.

verbose_json

duration is the billed audio length in seconds. segments and words are always empty.

{
  "task": "translate",
  "language": "en",
  "duration": 4.52,
  "text": "Hello, how are you?",
  "segments": [],
  "words": []
}

Tip: For transcription in the original language (not translated), use Speech to Text instead. For output modes like transliteration or code-mixing, use the mode parameter on the transcription endpoint.