# Whisper Diarization Whisper Large V3 Turbo delivers blazing-fast audio transcription with speaker diarization, converting conversations into accurate text with word- and sentence-level timestamps ## API Information - **Model Slug:** whisper-diarization - **Branded URL:** https://www.eachlabs.ai/openai/whisper/whisper-diarization - **Provider:** OpenAI - **Category:** Voice to Text - **Output Type:** object - **Status:** active - **Version:** 0.0.1 - **Base Cost:** Per-second pricing based on provider predict_time. Rate: $0.00108/sec from GPU tier. - **Estimated Processing Time:** 8 seconds - **Last Updated:** 2026-04-06 - **Interactive Demo:** https://www.eachlabs.ai/ai-models/whisper-diarization ## Pricing - **Charge Type:** dynamic - **Pricing Details:** Per-second pricing based on provider predict_time. Rate: $0.00108/sec from GPU tier. ### Pricing Rules | Condition | Pricing | | --- | --- | | Rule 1 | Per-second pricing based on provider predict_time. Rate: $0.00108/sec from GPU tier. | ## Input Schema | Parameter | Type | Required | Default | Constraints | Description | |-----------|------|----------|---------|-------------|-------------| | file_string | string | No | - | - | Either provide: Base64 encoded audio file, | | file_url | string | No | - | - | Or provide: A direct audio file URL | | file | string | Yes | - | - | Or an audio file | | group_segments | boolean | No | true | - | Group segments of same speaker shorter apart than 2 seconds | | num_speakers | integer | No | 2 | 1–50 | Number of speakers, leave empty to autodetect. | | translate | boolean | No | false | - | Translate the speech into English. | | language | string | No | en | - | Language of the spoken words as a language code like 'en'. Leave empty to auto detect language. | | prompt | string | No | - | - | Vocabulary: provide names, acronyms and loanwords in a list. Use punctuation for best accuracy. | ## Example Request ```bash curl -X POST https://api.eachlabs.ai/v1/prediction/ \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "whisper-diarization", "input": { "file": "https://storage.googleapis.com/magicpoint/inputs/whisper-diarization-input.mp3" } }' ``` ## Output Schema Response returned by `GET /v1/prediction/{id}` when the job completes: ```json { "status": "success", "predictionID": "string", "output": "object", "metrics": { "predict_time": "number (seconds)" } } ``` ## Polling ```bash curl https://api.eachlabs.ai/v1/prediction/{PREDICTION_ID} \ -H "Authorization: Bearer YOUR_API_KEY" ``` | Status | Meaning | |--------|---------| | `processing` | Still running — poll again | | `success` | Done — read `output` | | `error` | Failed — read `message` / `details` | ## Webhook (alternative to polling) Pass `"webhook_url": "https://your.host/path"` in the create request. Eachlabs POSTs this payload when the job ends: ```json { "exec_id": "prediction-uuid", "status": "succeeded", "output": "https://...", "error": "" } ``` `status` is `"succeeded"` or `"failed"`. `exec_id` equals the `predictionID` from create. Return 2xx within 30 seconds. ## Errors Error body: `{ "status": "error", "message": "...", "details": "..." }` | Code | Meaning | |------|---------| | `400` | Invalid input | | `401` | Missing / invalid `Authorization` bearer token | | `404` | Unknown model or prediction id | | `429` | Rate limit — 100 creates / min, 10 concurrent per key | | `5xx` | Retry with backoff | ## Overview **whisper-diarization — Voice-to-Text AI Model** whisper-diarization, powered by OpenAI's **Whisper Large V3 Turbo** architecture, revolutionizes voice-to-text transcription by delivering blazing-fast audio processing with integrated speaker diarization, accurately distinguishing speakers in conversations while providing word- and sentence-level timestamps. Developed as part of the Whisper family, this **voice-to-text AI model** tackles multi-speaker audio challenges that plague standard transcription tools, enabling precise conversion of meetings, podcasts, or interviews into searchable, timestamped text. With 6x faster inference than Whisper Large V3—thanks to its reduced 4 decoder layers and 809 million parameters—whisper-diarization maintains near-identical accuracy (within 1-2%) while handling 99+ languages. Ideal for developers seeking **OpenAI voice-to-text** solutions with diarization, it processes audio via simple API calls on Eachlabs, supporting chunking for large files. ## Usage Notes - API Base URL: `https://api.eachlabs.ai/v1` - Authentication: send `Authorization: Bearer YOUR_API_KEY`. Generate a key from the Eachlabs dashboard at https://www.eachlabs.ai/dashboard/api-keys. - File-typed parameters (`*_url`, `image_url`, `video_url`, `audio_url`, etc.) accept publicly-reachable HTTPS URLs only. Upload your asset first (GCS / S3 / your CDN) and pass the resulting URL. Data-URIs and localhost URLs are rejected. - For structured parameters (arrays / objects) send real JSON values, not stringified payloads. - Monetary values are reported in USD; per-token / per-megapixel rates may be billed in micro-cents internally. - Prefer `webhook_url` over polling for long-running predictions — see the Webhook Callback section.