Generate up to 20-second videos with native audio from a prompt, a start image, start/end frames, multiple keyframes, or a source clip. FLUX.3 chooses the matching workflow automatically from what you attach.
Added Aug 5, 2026
Approx. Price
$0.300 per video
Model Type
both
Settings
Generation controls available for this model.
Output Format
Default Duration
5
16 duration options
Aspect Ratio
Default
auto
Options (8)
Auto, Ultra-wide (21:9), Wide (2:1), Landscape (16:9) +4 more
Output frame shape; Auto chooses from the attached media or prompt.
Duration
Default
5
Options (16)
5 seconds, 6 seconds, 7 seconds, 8 seconds +12 more
Clip length from 5 to 20 seconds.
Generate Audio
Default
Yes
Generate synchronized audio and dialogue.
Render Quality
Default
full
Options (2)
Full quality, Draft preview
Draft creates a lower-cost preview; Full enables the highest output quality.
Resolution
Default
720p
Options (2)
720p, 1080p
1080p is available for full-quality generations.
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Examples
Loading examples…
Related video models
Compare FLUX.3 with similar models from the same provider or model family.
LTX-2.5 Fast
lightricks/ltx-2.5/fastSpeed-optimized audiovisual generation from text, an image, or a 2-20 second audio clip. Creates synchronized video and audio in one pass, with output up to 4K and optional start/end-frame control.
LTX-2.5 Pro
lightricks/ltx-2.5/proHigh-fidelity audiovisual generation from text, an image, or a 2-20 second audio clip. Creates polished synchronized video and audio in one pass, with 720p/1080p output and optional start/end-frame control.
Pixelcut Video Background Remover
pixelcut/video-background-removalRemove video backgrounds with frame-by-frame AI segmentation and temporally consistent edges. Supports transparent output, preset solid backgrounds, custom RGB backgrounds, and common video formats.
Bernini R Video
bernini-r-videoBernini R video generation and editing. A single NanoGPT model id routes prompt-only requests to text-to-video, up to 5 image inputs to reference-to-video, video inputs to edit-video, and image plus video inputs to reference-edit-video.
Luma Ray 3.2
luma/agent/ray/v3.2Cinematic text-to-video and image-to-video generation with strong motion control, optional reference images, seamless loops, 540p/720p/1080p output, and 5 or 10 second clips.
LTX-2.3 Quality
ltx-2.3-qualityLTX-2.3 Quality routes text, image, audio, reference video, extend-video, video-to-HDR, and optional LoRA inputs to FAL. Supports provider output sizes such as landscape, portrait, square, and auto with native synchronized audio.