# Bytedance | Omnihuman v1.5 ByteDance OmniHuman v1.5 generates expressive, full-body video performances from a single image and up to 30 seconds of audio with precise lip sync and emotionally responsive body language — at 1024×1024 resolution. Its dual System 1 / System 2 AI architecture understands the emotional and semantic meaning of audio before animating, producing cinematic-quality digital humans from minimal input. ## API Information - **Model Slug:** bytedance-omnihuman-v1-5 - **Branded URL:** https://www.eachlabs.ai/bytedance/omnihuman/bytedance-omnihuman-v1-5 - **Provider:** ByteDance - **Category:** Image to Video - **Output Type:** video - **Status:** active - **Version:** 0.0.1 - **Base Cost:** Per-second pricing based on generated duration: 0.12$/s - **Estimated Processing Time:** 140 seconds - **Last Updated:** 2026-07-22 - **Interactive Demo:** https://www.eachlabs.ai/ai-models/bytedance-omnihuman-v1-5 ## Pricing - **Charge Type:** dynamic - **Pricing Details:** Per-second pricing based on generated duration: 0.12$/s ### Pricing Rules | Condition | Pricing | | --- | --- | | Rule 1 | Per-second pricing based on generated duration: 0.12$/s | ## Input Schema | Parameter | Type | Required | Default | Constraints | Description | |-----------|------|----------|---------|-------------|-------------| | prompt | string | No | - | - | The text prompt used to guide the video generation. | | image_url | string | Yes | - | - | The URL of the image used to generate the video | | audio_url | string | Yes | - | - | The URL of the audio file to generate the video. Audio must be under 30s long for 1080p generation and under 60s long for 720p generation. | | turbo_mode | boolean | No | false | - | Generate a video at a faster rate with a slight quality trade-off. | | resolution | string | No | 1080p | 720p,1080p | The resolution of the generated video. Defaults to 1080p. 720p generation is faster and higher in quality. 1080p generation is limited to 30s audio and 720p generation is limited to 60s audio. | ## Example Request ```bash curl -X POST https://api.eachlabs.ai/v1/prediction/ \ -H "X-API-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "bytedance-omnihuman-v1-5", "input": { "audio_url": "https://storage.googleapis.com/magicpoint/inputs/bytedance-omnihuman-v1-5-input-audio.mp3", "image_url": "https://storage.googleapis.com/magicpoint/inputs/bytedance-omnihuman-v1-5-input-image.png" } }' ``` ## Output Schema Response returned by `GET /v1/prediction/{id}` when the job completes: ```json { "status": "success", "predictionID": "string", "output": "string (URL of generated video)", "metrics": { "predict_time": "number (seconds)" } } ``` ## Polling ```bash curl https://api.eachlabs.ai/v1/prediction/{PREDICTION_ID} \ -H "X-API-Key: YOUR_API_KEY" ``` | Status | Meaning | |--------|---------| | `processing` | Still running — poll again | | `success` | Done — read `output` | | `error` | Failed — read `message` / `details` | ## Webhook (alternative to polling) Pass `"webhook_url": "https://your.host/path"` in the create request. Eachlabs POSTs this payload when the job ends: ```json { "exec_id": "prediction-uuid", "status": "succeeded", "output": "https://...", "error": "" } ``` `status` is `"succeeded"` or `"failed"`. `exec_id` equals the `predictionID` from create. Return 2xx within 30 seconds. ## Errors Error body: `{ "status": "error", "message": "...", "details": "..." }` | Code | Meaning | |------|---------| | `400` | Invalid input | | `401` | Missing / invalid `X-API-Key` | | `404` | Unknown model or prediction id | | `429` | Rate limit — 100 creates / min, 10 concurrent per key | | `5xx` | Retry with backoff | ## Overview **bytedance-omnihuman-v1.5 — Image-to-Video AI Model** Bytedance's omnihuman-v1.5 transforms static images into expressive, full-body video performances by combining a reference image with audio input. Rather than generating video from text alone, this image-to-video AI model anchors the output to a specific person or character, enabling precise control over identity while the audio drives natural lip-sync and gesture animation. The model solves a critical problem for creators, marketers, and developers: producing film-grade talking-head and full-body videos without expensive studio setups or manual animation. Developed by Bytedance as part of the omnihuman family, bytedance-omnihuman-v1.5 specializes in audio-driven video synthesis with a focus on behavioral authenticity. The model generates videos up to 30 seconds in length, supporting HD resolution output at 1024×1024 pixels. Its core strength lies in synchronized lip-sync animation paired with emotionally responsive body language—the character doesn't just mouth words, but moves naturally in response to speech patterns and emotional tone embedded in the audio. ## Usage Notes - API Base URL: `https://api.eachlabs.ai/v1` - Authentication: send `X-API-Key: YOUR_API_KEY`. Generate a key from the Eachlabs dashboard at https://www.eachlabs.ai/dashboard/api-keys. - File-typed parameters (`*_url`, `image_url`, `video_url`, `audio_url`, etc.) accept publicly-reachable HTTPS URLs only. Upload your asset first (GCS / S3 / your CDN) and pass the resulting URL. Data-URIs and localhost URLs are rejected. - For structured parameters (arrays / objects) send real JSON values, not stringified payloads. - Monetary values are reported in USD; per-token / per-megapixel rates may be billed in micro-cents internally. - Prefer `webhook_url` over polling for long-running predictions — see the Webhook Callback section.