Audio-driven two-person avatar generation from a single image and left/right audio tracks. Supports simultaneous or sequential dialogue, 480p/720p output, and up to 30 seconds of audio.
Added May 23, 2026
Approx. Price
$0.150 per video
Model Type
image-to-video
Settings
Generation controls available for this model.
Output Format
Default Duration
N/A
Audio Order
Default
meanwhile
Options (3)
Together, Left then right, Right then left
Whether both speakers talk together or one after the other
Left Audio URL
Default
N/A
Public URL for the person on the left
Prompt
Default
N/A
Optional expression/style prompt
Resolution
Default
480p
Options (2)
480p, 720p
Video resolution
Right Audio URL
Default
N/A
Public URL for the person on the right
Benchmarks
Benchmarks
No benchmark data is available yet for this model.
Examples
Loading examples…
Related video models
Compare LongCat Avatar 1.5 Multi with similar models from the same provider or model family.
LongCat Avatar 1.5
wavespeed-ai/longcat-avatar-1.5Upgraded audio-driven talking or singing avatar generation from a single image with sharper lip sync and faster generation. Supports 480p/720p output up to 30 seconds.
LTX-2.3 Spicy Image-to-Video
wavespeed-ai/ltx-2.3-spicy/image-to-videoLTX-2.3 Spicy image-to-video turns a reference image and prompt into expressive clips with native audio, style-tuned motion, 480p/720p/1080p output, and 3-20 second durations.
LTX-2.3 Spicy LoRA Image-to-Video
wavespeed-ai/ltx-2.3-spicy/image-to-video-loraLTX-2.3 Spicy LoRA image-to-video adds selectable LoRA presets and per-LoRA strength overrides on top of the same reference-image animation workflow.
Music Video Generator
wavespeed-ai/music-video-generatorGenerate a lip-synced music video from an audio track, with optional reference portraits (1-3 images). Supports cinematic scene transitions up to 10 minutes at 480p or 720p.
LongCat Avatar
longcat-avatarAudio-driven talking or singing avatar generation from a single image with lip-synced motion and consistent identity. Supports 480p/720p output up to 2 minutes.
P-Video Avatar
pruna-ai/p-video/avatarImage-and-audio avatar video generation for speech-driven talking-head clips, with 720p and 1080p output.