Unified HappyHorse 1.0 video model. Routes to text-to-video, image-to-video, reference-to-video, or video edit based on the inputs you provide.
Added Apr 26, 2026
Approx. Price
$0.420 per video
Model Type
both
Settings
Generation controls available for this model.
Output Format
Default Duration
5
13 duration options
Aspect Ratio
Default
16:9
Options (5)
16:9 (Landscape), 9:16 (Portrait), 1:1 (Square), 4:3 +1 more
Used for text-only generations and optional for reference/edit flows.
Duration
Default
5
Options (13)
3 seconds, 4 seconds, 5 seconds, 6 seconds +9 more
3-15 seconds
Mode
Default
auto
Options (3)
Auto-detect, Edit Video, Reference to Video
Auto-detect uses your uploads. Multiple reference images route to reference-to-video. Video input routes to video edit.
Prompt Expansion
Default
true
Options (2)
true, false
Use AI to expand your prompt
Resolution
Default
720p
Options (2)
720p, 1080p
Output resolution
Benchmarks
Benchmarks
Human preference benchmarks sourced from Artificial Analysis.
Text to Video
#1 / 85
ELO
1290.0
Appearances
5,668
95% CI
-9/9
Image to Video
#7 / 79
ELO
1290.0
Appearances
5,710
95% CI
-10/10
Release Date 2026-04 · Matched as HappyHorse-1.0
Artificial Analysis APIRelated video models
Compare HappyHorse 1.0 with similar models from the same provider or model family.
HappyHorse 1.1
happyhorse-1.1Unified HappyHorse 1.1 video model. Routes to text-to-video, image-to-video, or reference-to-video with up to 9 reference images based on the inputs you provide.
LTX-2.5 Fast
lightricks/ltx-2.5/fastSpeed-optimized audiovisual generation from text, an image, or a 2-20 second audio clip. Creates synchronized video and audio in one pass, with output up to 4K and optional start/end-frame control.
LTX-2.5 Pro
lightricks/ltx-2.5/proHigh-fidelity audiovisual generation from text, an image, or a 2-20 second audio clip. Creates polished synchronized video and audio in one pass, with 720p/1080p output and optional start/end-frame control.
Wan 3.0 Image-to-Video
alibaba/wan-3.0/image-to-videoAnimate a first-frame image into a cinematic video with optional last-frame guidance, synchronized audio, deep-thinking controls, and 2–30 second output.
Wan 3.0 Reference-to-Video
alibaba/wan-3.0/reference-to-videoReference-guided video generation using images, videos, and audio for subject consistency, motion, timing, and scene continuity, with 2–30 second output.
Wan 3.0 Text-to-Video
alibaba/wan-3.0/text-to-videoCinematic text-to-video generation with synchronized audio, deep-thinking prompt interpretation, 2–30 second duration, and 480p, 720p, or 1080p output.