Hailuo S2V-01 (MiniMax): Character-Consistent AI Video from One Reference Image

Last verified: August 13, 2026

Status update: S2V-01 remains documented on MiniMax’s V1 Subject-Reference endpoint, while MiniMax-H3 is now the current general-purpose video model on Video Generation V2.

MiniMax Hailuo S2V-01 is an older, task-specific Subject Reference model for generating short videos from a face reference and a text prompt. MiniMax’s current V1 Subject-Reference reference still lists model: "S2V-01", so this page remains useful for existing integrations and searches that specifically target that workflow.

S2V-01 is not MiniMax’s current general-purpose video model. For a new integration, start by evaluating MiniMax-H3 through Video Generation V2. H3 accepts text plus supported image, video, and audio references and generates 4–15-second output at 768P or 2K with native stereo sound. H3’s generalized reference workflow is not automatically a drop-in replacement for S2V-01’s dedicated face-reference schema, so test identity retention before migrating an established S2V workflow.

Quick Answer: What Does MiniMax Hailuo S2V-01 Do?

MiniMax Hailuo S2V-01 creates AI videos from:

  • one subject reference image,
  • a text prompt describing the scene and motion,
  • and a Subject Reference workflow designed to preserve facial identity.

In MiniMax’s API documentation, Subject-Reference Video is described as a mode that uses a subject’s face photo plus a text description to generate a video while keeping the subject’s facial features consistent throughout the result.

The practical value is simple: if you want a recurring character, spokesperson, avatar, model, or brand persona to appear in multiple short video clips without completely changing face shape from shot to shot, S2V-01 is the Hailuo workflow built for that problem.

What Is MiniMax Hailuo S2V-01?

MiniMax Hailuo S2V-01 is the Subject Reference model in the Hailuo video ecosystem. MiniMax announced S2V-01 on January 10, 2025 as a model focused on maintaining consistent identity across video frames, even when the camera angle, pose, expression, lighting, or movement changes.

In the current MiniMax API reference, the Subject-Reference endpoint accepts model: "S2V-01", a required subject_reference array, and a text prompt of up to 2,000 characters. The same API reference also shows a prompt_optimizer setting that defaults to true, with the option to set it to false for more precise prompt control.

In plain English: S2V-01 is for “put this same face or character into this video scene.”

It is not the same as a standard image-to-video model, where the uploaded image usually acts as the first frame of the video. S2V-01 uses the uploaded image more like an identity reference, while the prompt controls the action and environment.

Where S2V-01 Fits After MiniMax H3

MiniMax’s current model catalog lists H3 as the current video model and places Hailuo 2.3, Hailuo 2.3 Fast, and Hailuo 02 under Legacy Models. S2V-01 remains separately documented on the V1 Subject-Reference endpoint. Documentation presence confirms the request schema; it does not guarantee access, pricing, or long-term availability for every account.

Model / workflowCurrent roleInput and outputBest useImportant boundary
MiniMax H3 / Video Generation V2Current general-purpose video modelText plus supported image, video, or audio context; 4–15s; 768P or 2K; native stereo soundNew multimodal video and reference workflowsPay-as-you-go V2 route; test face consistency rather than assuming S2V equivalence
S2V-01 / V1 Subject ReferenceOlder task-specific face-reference workflow that remains documentedSubject reference image plus promptExisting character-consistency integrationsVerify account access and billing; no clear S2V-specific price is published in the current central table
Hailuo 2.3 / 2.3 FastLegacy V1 video families2.3: text/image-to-video; Fast: image-to-videoMaintaining compatible older workflowsDo not describe either as MiniMax’s current video model
Hailuo 02Legacy V1 video modelText/image-to-videoExisting legacy integrationsNot the default recommendation for a new build

Choose S2V-01 when an existing workflow specifically depends on its face-reference schema. Evaluate H3 first for a new project, while measuring whether its generalized reference workflow preserves the exact identity behavior your use case requires.

S2V-01 vs Image-to-Video vs Text-to-Video

MiniMax’s older V1 Video Generation documentation separates text-to-video, image-to-video, first-and-last-frame, and Subject Reference workflows. H3’s V2 API uses a different content-array design covering text-to-video, first/last-frame image-to-video, and generalized reference-to-video. Keep the table below as an explanation of the S2V-01/V1 distinction, not as a map of the current H3 endpoint.

WorkflowWhat You ProvideWhat the Model Tries to ControlUse It When
Text-to-VideoA text promptThe entire scene from textYou do not need a specific person or exact image anchor.
Image-to-VideoA starting image + promptMotion evolving from the first frameYou want to animate a still image.
First-and-Last-Frame VideoStart image + end image + promptTransition between two framesYou need more control over beginning and ending composition.
S2V-01 Subject ReferenceFace/character reference + promptSubject identity across the generated clipYou need the same character to remain recognizable.

S2V-01 is not automatically better than image-to-video. It solves a different problem. If the uploaded image composition must remain intact, image-to-video may be more appropriate. If the face must remain recognizable while the scene changes, S2V-01 is the better fit.

Who Should Use MiniMax Hailuo S2V-01?

S2V-01 is a good fit for creators who need visual continuity across short AI video clips.

Best For

  • AI video creators building a recurring character.
  • Social media teams creating a consistent AI spokesperson.
  • Marketing teams using an approved model, mascot, or brand persona.
  • Indie filmmakers testing character shots before production.
  • Agencies making short concept clips for campaigns.
  • Developers integrating Subject Reference generation into internal tools.
  • Educators or presenters creating a consistent avatar-style presence.

Not Best For

  • Long narrative video production without editing.
  • Multi-character blocking where several identities must stay locked.
  • Product-only object reference workflows.
  • Highly precise scene control where prompt adherence matters more than identity.
  • Legal-risky uses involving celebrities, private individuals without consent, minors, or copyrighted characters.

MiniMax’s original S2V-01 announcement notes that the model may sometimes follow prompts less precisely than T2V or I2V and may show environmental morphing. That makes it useful, but not a guaranteed production shortcut for every scene.

How Subject Reference Works in Practice

The Subject Reference workflow uses a face image as the identity anchor and a prompt as the motion and scene instruction. In MiniMax’s API guide, the subject-reference example uses a subject_reference parameter with type: "character" and an image URL, then combines it with a prompt and model: "S2V-01".

A simplified request looks like this:

{
"model": "S2V-01",
"subject_reference": [
{
"type": "character",
"image": [
"https://example.com/reference-face.jpg"
]
}
],
"prompt": "A confident young founder walks through a modern studio, smiles at the camera, and gestures naturally while soft cinematic lighting reflects on the background.",
"prompt_optimizer": false
}

Use this only as a simplified structural example. Always check MiniMax’s current API documentation before production implementation, because model availability, parameters, limits, pricing, and account permissions can change. MiniMax’s video generation workflow is asynchronous: create a task, check task status, then retrieve the generated video file.

How to Create Character-Consistent AI Videos with S2V-01

Step 1: Choose the Right Reference Image

Your reference image is the foundation of the result. Use a clear image where the subject’s face is visible, well-lit, and not obscured.

A strong S2V-01 reference image usually has:

  • a single clear subject,
  • visible facial features,
  • neutral or natural expression,
  • no heavy filters,
  • no extreme blur,
  • no sunglasses or face-covering accessories,
  • enough resolution for facial detail,
  • simple background,
  • no confusing secondary faces.

Avoid images where the face is tiny, side-profile only, heavily stylized, or partially hidden. S2V-01 is designed around facial identity consistency, so the model needs a strong identity anchor.

Step 2: Decide the Clip’s Job Before Prompting

Before writing the prompt, define the purpose of the video:

  • Is it an ad variation?
  • A character test?
  • A talking-head style social clip?
  • A cinematic shot?
  • A fashion or lifestyle scene?
  • A brand mascot animation?
  • A product-supporting human scene?

A vague prompt creates vague motion. A strong prompt gives the model a specific action, setting, camera behavior, and visual style.

Step 3: Use a Prompt Formula

A practical S2V-01 prompt formula:

[Subject role] + [specific action] + [setting] + [camera movement] + [expression/body language] + [lighting/style] + [identity consistency instruction]

Example:

The same woman from the reference image walks slowly through a bright creative studio, holding a tablet and smiling confidently. The camera starts with a medium shot, then gently pushes in as she turns toward the lens. Natural daylight, clean commercial look, realistic facial features, stable identity, no exaggerated expression.

The phrase “stable identity” is not magic, but it can help keep your instruction focused. The real identity anchor is still the reference image.

Step 4: Keep Motion Specific but Not Overloaded

Short AI video models work best when the prompt describes one clear movement. Instead of asking for five actions in one clip, focus on a single beat.

Weak prompt:

Make her walk, dance, wave, speak, jump, turn around, and show a city in the background.

Stronger prompt:

The same woman from the reference image walks across a rainy city sidewalk, glances toward the camera, and gives a small confident smile. Slow handheld tracking shot, neon reflections, cinematic realism.

Step 5: Generate Variations, Then Select the Most Stable Clip

Expect to generate multiple versions. Character consistency can vary between outputs, especially with complex movement, wide-angle shots, fast camera changes, or dramatic lighting shifts.

For a production workflow, create several short variations, reject identity drift early, then edit the strongest clips together outside the model.

Reference Image Checklist

Use this checklist before uploading your subject image.

CheckWhy It Matters
Face clearly visibleS2V-01 needs enough identity information to preserve facial features.
One main subjectExtra faces can confuse the subject reference.
Good lightingShadows and blur reduce facial detail.
No heavy filtersFilters can become part of the identity and distort outputs.
Natural expressionExtreme expressions may carry into the generated video.
Commercial rights clearedYou should only use images you own or have permission to use.
No public figures or private people without consentLikeness and privacy risks vary by jurisdiction and use case.

Prompt Examples for S2V-01

1. Cinematic Founder Video

The same man from the reference image stands inside a modern startup office, looking thoughtful and confident. He walks slowly toward a glass wall, turns to the camera, and smiles. Subtle handheld camera movement, soft daylight, realistic cinematic style, stable facial identity.

2. Social Media Fashion Clip

The same woman from the reference image walks down a clean urban street wearing a stylish beige coat. The camera tracks her from the front as her hair moves naturally in the breeze. Bright lifestyle fashion look, shallow depth of field, natural smile, consistent face throughout.

3. Brand Mascot or Character Clip

The same illustrated character from the reference image stands in a colorful studio and gives an enthusiastic thumbs-up. The camera gently zooms in while the character smiles. Clean animated commercial style, consistent character design, no face changes.

4. Professional Spokesperson Shot

The same person from the reference image sits at a desk in a minimal studio, looking into the camera with a calm, trustworthy expression. The camera remains mostly static with a slight push-in. Soft key light, professional corporate tone, realistic facial consistency.

5. Product-Supporting Lifestyle Scene

The same woman from the reference image holds a reusable water bottle while walking through a sunlit park. She looks at the bottle, then smiles toward the camera. Natural lifestyle advertising style, warm morning light, stable identity, smooth motion.

Troubleshooting: Common S2V-01 Problems and Fixes

ProblemLikely CausePractical Fix
Face changes mid-videoReference image is weak, motion is too extreme, or camera angle changes too aggressivelyUse a clearer reference, reduce motion, request a medium shot, and avoid fast rotations.
Subject looks similar but not identicalPrompt asks for style changes that override identityRemove heavy style terms, keep lighting natural, and describe the subject more consistently.
Environment morphsScene is too complex or prompt is overloadedSimplify the background and ask for one clear setting.
Prompt not followed exactlyS2V-01 prioritizes subject consistency and may be less precise than T2V/I2VShorten the prompt and focus on one action. MiniMax itself notes that S2V-01 may sometimes follow prompts less precisely than T2V or I2V.
Unnatural facial expressionPrompt over-specifies emotion or reference image has an extreme expressionUse a neutral reference and prompt for subtle expressions.
Subject becomes too stylizedPrompt includes strong art-style termsRemove style-heavy words and use “realistic,” “natural lighting,” or “commercial video style.”
Poor output for multiple peopleS2V-01 was introduced around single-subject identity consistency, and MiniMax described multi-subject references as an area for future improvementUse one subject per clip or verify whether current tools support your exact multi-subject workflow.

Access, API Notes, Pricing, and Limits

As of August 13, 2026, MiniMax’s Subject-Reference reference still documents POST /v1/video_generation with model: "S2V-01", a required subject_reference array, and a prompt of up to 2,000 characters. Do not transfer H3 parameters to that endpoint: H3 uses POST /v2/video_generation and a content array.

MiniMax’s current rate-limit table separates the routes. The V1 Video Generation row for the Hailuo series lists 5 RPM for free-tier accounts and 20 RPM for paid-tier accounts. Video Generation V2 for MiniMax-H3 instead lists a maximum of 2 concurrent tasks on the free tier and 15 on the paid tier. RPM and CONN are not interchangeable.

The current pay-as-you-go page lists H3 output and input-material pricing and places Hailuo 2.3, 2.3 Fast, and Hailuo 02 under Legacy Models. It still does not publish a clear S2V-01-specific price. Verify S2V-01 access and cost in the MiniMax console or account documentation before production use.

Third-party platforms may provide S2V-01 access under their own pricing and terms. For example, fal.ai lists a MiniMax Video-01 Subject Reference model page and displays a per-video cost on that page, while Replicate lists MiniMax Video-01 with a subject reference option for S2V-01. Treat those as third-party access routes, not as official MiniMax pricing.

If you are using the Hailuo web app rather than the API, check your current account plan and terms before commercial use. Hailuo’s subscription terms distinguish free-user watermarked downloads from paid-user no-watermark downloads and state that, for content generated and downloaded while on a paid subscription plan, Hailuo does not claim ownership and the user retains intellectual property rights, including commercial use. That should still be read alongside Hailuo’s Terms of Service, content standards, privacy policy, and any account-specific restrictions.

Practical Production Workflow

A safe production workflow for S2V-01 looks like this:

  1. Start with an approved reference image.
    Confirm you own the image or have permission to use it.
  2. Create a one-sentence creative brief.
    Example: “Generate a consistent AI spokesperson walking through a clean studio for a SaaS launch teaser.”
  3. Write three prompt variants.
    Keep each variant focused on one action and one environment.
  4. Generate multiple outputs.
    Do not expect the first generation to be final.
  5. Reject identity drift early.
    If the face changes noticeably, fix the reference image or simplify the prompt.
  6. Edit outside the model.
    Use a video editor for pacing, captions, music, brand graphics, and final color.
  7. Review rights and platform terms.
    Confirm watermark, commercial-use, consent, and privacy requirements before publishing.

When to Keep S2V-01 and When to Start With H3

Keep S2V-01 when an existing integration specifically needs the documented subject_reference face workflow, the account still accepts model: "S2V-01", and the team has verified its actual resolution, duration, cost, and output quality.

Start with H3 for a new integration when you need the current model, Video Generation V2, 4–15-second output, 768P or 2K, native stereo sound, first/last-frame control, or multimodal image, video, and audio references. H3 is broader, but benchmark its generalized reference workflow before claiming the same face-locking behavior as S2V-01.

Use Hailuo 2.3, Hailuo 2.3 Fast, or Hailuo 02 only when maintaining a compatible legacy workflow or when account-specific testing gives a documented reason to retain them. Do not present them as the current default for new MiniMax video work.

Safety, Consent, Copyright, and Privacy

S2V-01 is powerful because it can preserve a person’s likeness. That is also why it requires careful use.

Do not upload or generate videos using someone’s face unless you have the right to do so. This is especially important for private individuals, employees, customers, minors, public figures, celebrities, and anyone whose likeness could be used misleadingly.

Hailuo’s Terms of Service say user contributions and generated content must comply with content standards, must not infringe intellectual property or other rights, and must not violate rights of publicity or privacy. The same terms state that users are responsible for the legality, reliability, accuracy, and appropriateness of submitted and generated content.

Privacy also matters because face-reference workflows involve facial images. Hailuo’s privacy policy says certain mobile app features may process uploaded face photos and temporarily derive facial feature data or landmarks for video generation, and states that this facial data is used for the processing task rather than identification or authentication.

There is also a broader copyright context around Hailuo. Reuters reported in May 2026 that a U.S. federal judge denied MiniMax’s motion to dismiss a copyright lawsuit brought by Disney, Universal, and Warner Bros. Discovery over allegations involving Hailuo AI; the case was allowed to proceed, not finally decided.

Practical rule: use S2V-01 with assets you own, licensed characters, approved brand mascots, consenting talent, or original fictional characters. Do not use it to imitate protected characters, celebrities, private people, or anyone who has not given permission.

Pros and Cons of MiniMax Hailuo S2V-01

ProsCons
Designed specifically for subject and facial identity consistencyMay follow prompts less precisely than T2V or I2V in some cases.
Needs only one reference image for the subjectNot ideal for complex multi-subject scenes unless current support is verified.
Useful for recurring characters, AI spokespeople, and brand personasOutputs still need human review for drift, artifacts, and rights issues.
Available through MiniMax’s Subject-Reference API path as of the latest source checkOfficial S2V-specific pricing was not clearly listed in the checked MiniMax video package table.
Can be integrated into developer workflows through asynchronous API generationThird-party platform pricing and terms may differ from MiniMax’s official platform.

Best Use Cases

AI Spokesperson Clips

S2V-01 can help create short, consistent clips of a brand spokesperson or fictional persona. Keep the shot simple: medium framing, subtle gestures, stable lighting, and one clear message.

Brand Mascot Video

If your brand has a licensed or original mascot, S2V-01 can help animate it across short scenes. Use consistent descriptions of the mascot’s visual traits in every prompt.

Character Tests for Film and Storyboarding

For previsualization, S2V-01 can help test how a character might look in different settings. Use it for concepting, not as a replacement for final production review.

Social Ad Variations

You can generate several short variations of a model or character in different locations, then edit the best outputs into a campaign test.

Creator Avatars

Creators can use an approved self-image to explore avatar-style clips for intros, teasers, and visual hooks. Make sure the final result is not misleading if the clip implies real behavior or endorsement.

When Not to Use S2V-01

Do not use S2V-01 when:

  • you need a long, polished final video from one generation,
  • you need exact choreography,
  • you need multiple named characters to remain stable,
  • the face reference is low quality,
  • you do not have permission to use the person’s image,
  • the character is copyrighted and not licensed,
  • the output could mislead viewers into thinking a real person said or did something they did not.

For professional publishing, add disclosure where appropriate and keep records of consent, licensing, prompts, and source assets.

FAQ

What does S2V mean in MiniMax Hailuo S2V-01?

MiniMax officially calls S2V-01 a Subject Reference model. Some creators and third-party platforms use “subject-to-video” as a practical shorthand, but the official MiniMax API category is Subject-Reference to Video Generation. In practice, the subject reference image anchors identity while the prompt controls the generated video scene.

Is S2V-01 the same as normal image-to-video?

No. Normal image-to-video usually animates a starting image. S2V-01 uses a subject reference image, especially a face photo, to preserve identity while generating a scene from a prompt. MiniMax’s video generation guide separates image-to-video from Subject Reference as different modes.

Can MiniMax Hailuo S2V-01 create a video from one image?

Yes, the workflow is designed around using a subject reference image and a prompt. MiniMax’s S2V-01 announcement says the model can generate character-consistent videos from one reference image, and the API reference requires subject_reference images for the S2V-01 model.

Does S2V-01 support objects as references?

The checked official MiniMax sources emphasize faces, characters, and facial identity. MiniMax’s original announcement described object references and multi-subject references as future improvement areas, so do not assume robust object-reference support unless your current MiniMax or platform account explicitly confirms it.

Does S2V-01 support multiple subjects?

Do not assume reliable multi-subject identity locking. MiniMax’s S2V-01 announcement said future improvements would include multi-subject references, which implies the initial focus was single-subject consistency. Verify current support before building a workflow around multiple people.

Is S2V-01 still available?

It remains documented on MiniMax’s V1 Subject-Reference API as of August 13, 2026, but H3 is the current general-purpose video model. Documentation does not guarantee entitlement for every account, region, or future date, so verify a test request and current billing before production use.

Is S2V-01 free?

Do not assume it is free. The official sources checked did not provide a clear S2V-01-specific point deduction table. Hailuo’s subscription terms discuss free and paid download conditions, while third-party providers may set their own per-video pricing. Always check the current pricing page inside the platform you use.

What resolution and duration does S2V-01 support?

MiniMax’s V1 Subject-Reference guide shows an example using model: "S2V-01", duration: 6, and resolution: "1080P". Verify the accepted values in the official reference and your account before production use. Do not apply H3’s 4–15-second, 768P/2K limits to S2V-01; they belong to the separate V2 endpoint.

How do I reduce identity drift?

Use a clear face reference, avoid extreme camera rotations, keep the prompt focused on one action, use medium or close framing, avoid heavy stylization, and generate multiple variations. If the face changes, simplify the scene before trying more complex movement.

Can I use a real person’s photo?

Only use a real person’s photo if you have the right and consent to do so. Hailuo’s terms say user-generated content must not violate rights of privacy or publicity and that users are responsible for generated content.

Can I use copyrighted characters?

Do not use copyrighted characters unless you have the necessary rights. Hailuo’s terms prohibit content that infringes copyright, trademark, or other rights, and Reuters reported that copyright claims involving Hailuo AI and major entertainment studios are currently proceeding in U.S. court.

More on Hailuo: see the Hailuo AI hub for every model compared side by side, or Is Hailuo AI free? for what each access route actually costs.