# Whisper Whisper is designed to turn speech into text across multiple languages. ## API Information - **Model Slug:** whisper - **Branded URL:** https://www.eachlabs.ai/openai/whisper/whisper - **Provider:** OpenAI - **Category:** Voice to Text - **Output Type:** text - **Status:** active - **Version:** 0.0.1 - **Base Cost:** Execution-time pricing: $0.001265/sec based on run_time - **Estimated Processing Time:** 8 seconds - **Last Updated:** 2026-04-16 - **Interactive Demo:** https://www.eachlabs.ai/ai-models/whisper ## Pricing - **Charge Type:** dynamic - **Pricing Details:** Execution-time pricing: $0.001265/sec based on run_time ### Pricing Rules | Condition | Pricing | | --- | --- | | Rule 1 | Execution-time pricing: $0.001265/sec based on run_time | ## Input Schema | Parameter | Type | Required | Default | Constraints | Description | |-----------|------|----------|---------|-------------|-------------| | audio_url | string | Yes | - | - | URL of the audio file to transcribe. Supported formats: mp3, mp4, mpeg, mpga, m4a, wav or webm. | | task | string | No | transcribe | transcribe,translate | Task to perform on the audio file. Either transcribe or translate. | | language | string | No | - | af,am,ar,as,az,ba,be,bg,bn,bo,br,bs,ca,cs,cy,da,de,el,en,es,et,eu,fa,fi,fo,fr,gl,gu,ha,haw,he,hi,hr,ht,hu,hy,id,is,it,ja,jw,ka,kk,km,kn,ko,la,lb,ln,lo,lt,lv,mg,mi,mk,ml,mn,mr,ms,mt,my,ne,nl,nn,no,oc,pa,pl,ps,pt,ro,ru,sa,sd,si,sk,sl,sn,so,sq,sr,su,sv,sw,ta,te,tg,th,tk,tl,tr,tt,uk,ur,uz,vi,yi,yo,zh | Language of the audio file. If set to null, the language will be automatically detected. Defaults to null. If translate is selected as the task, the audio will be translated to English, regardless of the language selected. | | diarize | boolean | No | false | - | Whether to diarize the audio file. Defaults to false. Setting to true will add costs proportional to diarization inference time. | | chunk_level | string | No | segment | none,segment,word | Level of the chunks to return. Either none, segment or word. `none` would imply that all of the audio will be transcribed without the timestamp tokens, we suggest to switch to `none` if you are not satisfied with the transcription quality, since it will usually improve the quality of the results. Switching to `none` will also provide minor speed ups in the transcription due to less amount of generated tokens. Notice that setting to none will produce **a single chunk with the whole transcription**. | | version | string | No | 3 | 3 | Version of the model to use. All of the models are the Whisper large variant. | | batch_size | integer | No | 64 | 1–64 | - | | prompt | string | No | - | - | Prompt to use for generation. Defaults to an empty string. | | num_speakers | integer | No | - | 1–0 | Number of speakers in the audio file. Defaults to null. If not provided, the number of speakers will be automatically detected. | ## Example Request ```bash curl -X POST https://api.eachlabs.ai/v1/prediction/ \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "whisper", "input": { "audio_url": "https://storage.googleapis.com/magicpoint/inputs/wizper-input.wav" } }' ``` ## Output Schema Response returned by `GET /v1/prediction/{id}` when the job completes: ```json { "status": "success", "predictionID": "string", "output": "text", "metrics": { "predict_time": "number (seconds)" } } ``` ## Polling ```bash curl https://api.eachlabs.ai/v1/prediction/{PREDICTION_ID} \ -H "Authorization: Bearer YOUR_API_KEY" ``` | Status | Meaning | |--------|---------| | `processing` | Still running — poll again | | `success` | Done — read `output` | | `error` | Failed — read `message` / `details` | ## Webhook (alternative to polling) Pass `"webhook_url": "https://your.host/path"` in the create request. Eachlabs POSTs this payload when the job ends: ```json { "exec_id": "prediction-uuid", "status": "succeeded", "output": "https://...", "error": "" } ``` `status` is `"succeeded"` or `"failed"`. `exec_id` equals the `predictionID` from create. Return 2xx within 30 seconds. ## Errors Error body: `{ "status": "error", "message": "...", "details": "..." }` | Code | Meaning | |------|---------| | `400` | Invalid input | | `401` | Missing / invalid `Authorization` bearer token | | `404` | Unknown model or prediction id | | `429` | Rate limit — 100 creates / min, 10 concurrent per key | | `5xx` | Retry with backoff | ## Overview **whisper — Voice-to-Text AI Model** Whisper transforms spoken audio into accurate text transcripts across nearly 100 languages, solving the challenge of reliable multilingual speech recognition for developers, creators, and businesses seeking robust **voice-to-text AI models**. Developed by OpenAI as part of the Whisper family, this open-source model excels in handling diverse accents, background noise, and real-world audio conditions where traditional systems falter. Trained on 680,000 hours of labeled data, Whisper delivers 95%+ accuracy in optimal settings, making it a go-to for **OpenAI voice-to-text** applications like transcription APIs and live captioning. ## Usage Notes - API Base URL: `https://api.eachlabs.ai/v1` - Authentication: send `Authorization: Bearer YOUR_API_KEY`. Generate a key from the Eachlabs dashboard at https://www.eachlabs.ai/dashboard/api-keys. - File-typed parameters (`*_url`, `image_url`, `video_url`, `audio_url`, etc.) accept publicly-reachable HTTPS URLs only. Upload your asset first (GCS / S3 / your CDN) and pass the resulting URL. Data-URIs and localhost URLs are rejected. - For structured parameters (arrays / objects) send real JSON values, not stringified payloads. - Monetary values are reported in USD; per-token / per-megapixel rates may be billed in micro-cents internally. - Prefer `webhook_url` over polling for long-running predictions — see the Webhook Callback section.