LTX-2.3 Spicy image-to-video turns a reference image and prompt into expressive clips with native audio, style-tuned motion, 480p/720p/1080p output, and 3-20 second durations.
Added Jun 18, 2026
Approx. Price
$0.100 per video
Model Type
image-to-video
Settings
Generation controls available for this model.
Output Format
Default Duration
5
18 duration options
Duration
Default
5
Options (18)
3 seconds, 4 seconds, 5 seconds, 6 seconds +14 more
3 to 20 seconds. Billing has a 5 second minimum.
Resolution
Default
480p
Options (3)
480p, 720p, 1080p
Output resolution
Seed
Default
-1
Use -1 for random output or set a seed for repeatable variations.
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Related video models
Compare LTX-2.3 Spicy Image-to-Video with similar models from the same provider or model family.
LTX-2 19B
ltx-2-19bUnified LTX-2 19B model for text-to-video or image-to-video with synchronized audio. Supports optional LoRA adapters for custom styles.
Lightricks LTX-2 Fast
lightricks-ltx-2-fastHigh-speed LTX-2 pipeline tuned for rapid iterations. Convert text or a single image into cinematic clips with synchronized audio in seconds.
Lightricks LTX-2 Pro
lightricks-ltx-2-proFlagship LTX-2 stack for production-ready motion. Generates synchronized audio and rich camera moves from text prompts or reference images.
LTX-2.3 Spicy LoRA Image-to-Video
wavespeed-ai/ltx-2.3-spicy/image-to-video-loraLTX-2.3 Spicy LoRA image-to-video adds selectable LoRA presets and per-LoRA strength overrides on top of the same reference-image animation workflow.
LongCat Avatar 1.5
wavespeed-ai/longcat-avatar-1.5Upgraded audio-driven talking or singing avatar generation from a single image with sharper lip sync and faster generation. Supports 480p/720p output up to 30 seconds.
LongCat Avatar 1.5 Multi
wavespeed-ai/longcat-avatar-1.5/multiAudio-driven two-person avatar generation from a single image and left/right audio tracks. Supports simultaneous or sequential dialogue, 480p/720p output, and up to 30 seconds of audio.