# Wizper with Timestamp Wizper with Timestamp is a multilingual speech recognition and translation model built on Whisper v3 that transcribes audio with precise word-level timestamps. It delivers fast, accurate, and time-aligned transcripts, making it ideal for subtitles, media indexing, and real-time transcription workflows ## API Information - **Model Slug:** wizper-with-timestamp - **Branded URL:** https://www.eachlabs.ai/openai/whisper/wizper-with-timestamp - **Provider:** Openai - **Category:** Voice to Text - **Output Type:** object - **Status:** active - **Version:** 0.0.1 - **Base Cost:** $0.00108 per processing second - **Estimated Processing Time:** 1 seconds - **Last Updated:** 2026-08-16 - **Interactive Demo:** https://www.eachlabs.ai/ai-models/wizper-with-timestamp ## Pricing - **Charge Type:** dynamic - **Estimate:** $0.00108 per processing second - **Pricing Details:** default: $0.00108 ## Input Schema | Parameter | Type | Required | Default | Constraints | Description | |-----------|------|----------|---------|-------------|-------------| | audio_url | string | Yes | - | - | URL of the audio file to transcribe. Supported formats: mp3, mp4, mpeg, mpga, m4a, wav or webm. | | task | string | No | transcribe | transcribe, translate | Task to perform on the audio file. Either transcribe or translate. | | language | string | No | - | af,am,ar,as,az,ba,be,bg,bn,bo,br,bs,ca,cs,cy,da,de,el,en,es,et,eu,fa,fi,fo,fr,gl,gu,ha,haw,he,hi,hr,ht,hu,hy,id,is,it,ja,jw,ka,kk,km,kn,ko,la,lb,ln,lo,lt,lv,mg,mi,mk,ml,mn,mr,ms,mt,my,ne,nl,nn,no,oc,pa,pl,ps,pt,ro,ru,sa,sd,si,sk,sl,sn,so,sq,sr,su,sv,sw,ta,te,tg,th,tk,tl,tr,tt,uk,ur,uz,vi,yi,yo,zh | Language of the audio file. If translate is selected as the task, the audio will be translated to English, regardless of the language selected. If `None` is passed, the language will be automatically detected. This will also increase the inference time. | | chunk_level | string | No | segment | - | Level of the chunks to return. | | max_segment_len | integer | No | 29 | 10–29 | Maximum speech segment duration in seconds before splitting. | | merge_chunks | boolean | No | true | - | Whether to merge consecutive chunks. When enabled, chunks are merged if their combined duration does not exceed max_segment_len. | | version | string | No | 3 | - | Version of the model to use. All of the models are the Whisper large variant. | ## Example Request ```bash curl -X POST https://api.eachlabs.ai/v1/prediction/ \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "wizper-with-timestamp", "input": { "audio_url": "https://cdn-us.eachlabs.ai/defaults/a8c3c8b8f7f893180f019703d657b7e5dec788d068ebe320b1eacc8db006d7ee.wav" } }' ``` ## Output Schema Response returned by `GET /v1/prediction/{id}` when the job completes: ```json { "status": "success", "predictionID": "string", "output": "object", "metrics": { "predict_time": "number (seconds)" } } ``` ## Polling ```bash curl https://api.eachlabs.ai/v1/prediction/{PREDICTION_ID} \ -H "Authorization: Bearer YOUR_API_KEY" ``` | Status | Meaning | |--------|---------| | `processing` | Still running — poll again | | `success` | Done — read `output` | | `error` | Failed — read `message` / `details` | ## Webhook (alternative to polling) Pass `"webhook_url": "https://your.host/path"` in the create request. Eachlabs POSTs this payload when the job ends: ```json { "exec_id": "prediction-uuid", "status": "succeeded", "output": "https://...", "error": "" } ``` `status` is `"succeeded"` or `"failed"`. `exec_id` equals the `predictionID` from create. Return 2xx within 30 seconds. ## Errors Error body: `{ "status": "error", "message": "...", "details": "..." }` | Code | Meaning | |------|---------| | `400` | Invalid input | | `401` | Missing / invalid `Authorization` bearer token | | `404` | Unknown model or prediction id | | `429` | Rate limit — 100 creates / min, 10 concurrent per key | | `5xx` | Retry with backoff | ## Overview **wizper-with-timestamp — Voice-to-Text AI Model** wizper-with-timestamp, developed by OpenAI as part of the **Whisper** family, delivers multilingual speech recognition with precise word-level timestamps, enabling accurate time-aligned transcripts for audio files. Built on Whisper v3, this **voice-to-text AI model** excels in transcribing long-form audio like videos or meetings, outputting clean text with timestamps ideal for subtitles and media indexing. Developers seeking **OpenAI voice-to-text** solutions with timestamp precision find wizper-with-timestamp perfect for real-time workflows and batch processing, supporting formats like mono WAV at 16kHz for optimal performance. ## Usage Notes - API Base URL: `https://api.eachlabs.ai/v1` - Authentication: send `Authorization: Bearer YOUR_API_KEY`. Generate a key from the Eachlabs dashboard at https://www.eachlabs.ai/dashboard/api-keys. - File-typed parameters (`*_url`, `image_url`, `video_url`, `audio_url`, etc.) accept publicly-reachable HTTPS URLs only. Upload your asset first (GCS / S3 / your CDN) and pass the resulting URL. Data-URIs and localhost URLs are rejected. - For structured parameters (arrays / objects) send real JSON values, not stringified payloads. - Monetary values are reported in USD; per-token / per-megapixel rates may be billed in micro-cents internally. - Prefer `webhook_url` over polling for long-running predictions — see the Webhook Callback section.