HeyGen Video 1.0 is a single general-purpose video generation model (heygen-video-1) built on MiniMax H3 and post-trained by HeyGen that produces 5-15 second clips with integrated audio from a single call. It supports three modes - text_to_video (prompt only), image_to_video (prompt plus a single image used as the literal first frame) and reference_to_video (prompt plus up to 12 references: images, videos, audio) - and is accessed via POST /v3/models/videos with retrieval by GET /v3/models/videos/{video_id}. Requests include fields like duration, resolution, aspect_ratio and seed (for deterministic results). Submissions return a pending/processing/completed/failed/cancelled lifecycle; successful jobs return a signed video_url. The API supports idempotency keys and an optional HTTPS callback_url/callback_id to receive a single terminal event, but polling is recommended as a fallback.
The model’s strengths and usage rules focus on predictable, physically consistent short shots: objects hold form and position, audio matches the described space, and it follows prompts literally. It excels at one subject/place/action with a static or slightly moving camera, and is weakest on long on-screen text, soft organic motion, and close hand work. Prompt guidance emphasizes long, detailed prompts; specifying camera/lens/aperture; forbidding beauty/retouching and invented branding; spelling any in-shot lettering exactly; describing sound and contact points; and using a fixed seed to iterate. Reference_to_video assigns ordinal labels to supplied assets (first image = <image_0>, etc.), with maximums of 9 images, 3 videos and 3 audio files (12 total).
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.