black-forest-labs/flux-3
Generate video with synchronized audio from text, images, or video. FLUX 3 is Black Forest Labs' multimodal model (early access preview).
Capabilities
Cost
Community model (estimated from hardware time)
Input Parameters
| Name | Type | Description | Default | Constraints |
|---|---|---|---|---|
prompt* | string | Text description of the video to generate. Plain language works — the prompt is interpreted and expanded before generation. Describe the scene, action, camera moves, and any audio you want. | — | — |
aspect_ratio | string | Aspect ratio of the generated video. 'auto' picks a ratio from your prompt and any inputs. | "auto" | auto21:92:116:94:31:13:49:16 |
draft | boolean | Generate a fast, low-cost draft preview instead of a full-quality clip. Drafts are 720p. | false | — |
duration | string | Length of the generated clip in seconds. 'auto' lets the model pick to fit the content. Three or more images need an explicit duration. | "auto" | auto567891011121314151617181920 |
generate_audio | boolean | Generate synchronized audio (ambient sound, speech, effects). Set to false for a silent clip. | true | — |
images | array | Optional images to drive the video, placed on screen pixel for pixel. One image opens the clip (image-to-video). Two images start and end it. Three or more (up to 10) become a storyboard — the first starts it, the last ends it, and the rest fall evenly in between (needs a duration). Leave empty for text-to-video. Must be PNG, JPEG, or WebP. | | — |
resolution | string | Output resolution. | "720p" | 720p1080p |
safety_tolerance | integer | Moderation tolerance for input and output, 0 (strictest) to 4 (most permissive). Requests with image or video inputs are limited to 2. | 2 | min: 0, max: 4 |
start_video | string(uri) | A video to continue from its final frames. Use this to extend a shot or chain generations into a longer sequence. Must be an mp4, at most 50MB and 15 seconds. Can't be combined with images. | — | — |
promptrequiredstringText description of the video to generate. Plain language works — the prompt is interpreted and expanded before generation. Describe the scene, action, camera moves, and any audio you want.
aspect_ratiostringAspect ratio of the generated video. 'auto' picks a ratio from your prompt and any inputs.
"auto"draftbooleanGenerate a fast, low-cost draft preview instead of a full-quality clip. Drafts are 720p.
falsedurationstringLength of the generated clip in seconds. 'auto' lets the model pick to fit the content. Three or more images need an explicit duration.
"auto"generate_audiobooleanGenerate synchronized audio (ambient sound, speech, effects). Set to false for a silent clip.
trueimagesarrayOptional images to drive the video, placed on screen pixel for pixel. One image opens the clip (image-to-video). Two images start and end it. Three or more (up to 10) become a storyboard — the first starts it, the last ends it, and the rest fall evenly in between (needs a duration). Leave empty for text-to-video. Must be PNG, JPEG, or WebP.
resolutionstringOutput resolution.
"720p"safety_toleranceintegerModeration tolerance for input and output, 0 (strictest) to 4 (most permissive). Requests with image or video inputs are limited to 2.
2min: 0, max: 4start_videostringA video to continue from its final frames. Use this to extend a shot or chain generations into a longer sequence. Must be an mp4, at most 50MB and 15 seconds. Can't be combined with images.
93515f53aad7Updated: 9/20/202627.8K runs
cinemasetfree