MiniMax H3: 2K Video, Native Stereo Audio, Open Weights, Pricing & Limits

Verified MiniMax H3 guide covering 2K video, 4–15-second clips, native stereo audio, multimodal inputs, API pricing, open-weight limits, and Hailuo comparisons.

Last verified: September 30, 2026. Subscription coverage was rechecked against MiniMax’s new M Plan page on that date. H3 prices, output limits, input limits, rate limits, and access routes were last rechecked against MiniMax’s live documentation on September 24, 2026, when a comparison with MiniMax H3 Max was added. License and open-weight facts were last verified August 11, 2026.

Quick answer: MiniMax H3 is MiniMax’s general-purpose video-and-audio generation system and one of its two current video models; the other is the speed-focused MiniMax H3 Max. The hosted H3 service accepts text, images, video, and audio as context, then generates clips from 4 to 15 seconds at 768P or 2K with native 32 kHz stereo sound. The exact API model ID is MiniMax-H3. MiniMax has also released H3-Base weights for local 768P generation, but the full 2K workflow is not completely open: H3-Context-IR remains hosted and H3-Regenerate-2K has not yet been released as open weights.

Independent-site notice: MiniMax-AI.chat is an independent educational website. It is not MiniMax’s official website and is not endorsed by or affiliated with MiniMax. Product names and trademarks belong to their respective owners. Access, pricing, licensing, moderation, availability, and provider terms are controlled by MiniMax.

Evidence status: Specifications, prices, access routes, examples, and model claims on this page were checked against MiniMax’s official announcement, API documentation, GitHub repository, Hugging Face model card, and license. The linked example videos were produced and published by MiniMax; they are not independent benchmark results from this site. We have not yet completed a paid, repeatable H3 benchmark, so visual-quality claims are identified as provider claims rather than our findings.

Bottom line: H3 is a major architectural and workflow change from Hailuo 2.3, not a simple version-number update. It adds native synchronized audio, longer and higher-resolution output, richer multimodal references, a new V2 API, and publicly released base weights. Use the hosted API for the simplest complete 2K workflow; treat local H3-Base as a 768P deployment with a restricted community license.

On this page

MiniMax H3 at a glance

SpecificationVerified value
Official release dateJuly 31, 2026
Hosted API model IDMiniMax-H3
Current roleOne of MiniMax’s two current video models, alongside H3 Max; H3 is the general-purpose, 2K-capable option
Hosted outputVideo with native stereo audio
Hosted resolutions768P and 2K
Output duration4-15 seconds, in whole-second values
Frame rate24 FPS
Audio32 kHz stereo
Documented dialogue languages with stable supportArabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish
Hosted inputsText, images, video, and audio
API endpointPOST https://api.minimax.io/v2/video_generation
PAYG output price$0.08/second at 768P; $0.13/second at 2K
Open-weights statusH3-Base FL2VA and Ref2VA checkpoints released in BF16 under the MiniMax H3 Community License
Local open-weight output768P; the complete local 2K stack is not yet open

The terms 768P and 2K above are the official API labels. MiniMax does not define 2K on the reviewed pages as one universal width-by-height pair, so this page does not convert it into an unsupported pixel-dimension promise.

What is MiniMax H3?

MiniMax H3 is a general-purpose multimodal generation system announced by MiniMax on July 31, 2026. It uses natural-language instructions to describe how text, images, video, and audio references relate to the requested output. Rather than treating text-to-video, image-to-video, motion reference, sound effects, dialogue, music, and editing as completely separate products, H3 is designed to handle them within one broader context-and-generation workflow.

The official H3 announcement says the model was built for instruction following, text and brand rendering, motion transfer, multimodal editing, and commercial content creation. These are MiniMax’s release claims. They should not be read as independent findings until the same prompts, references, costs, failure rates, and review criteria have been tested across competing models.

How does MiniMax H3 relate to Hailuo?

Hailuo is both a product name and part of MiniMax’s earlier video-model naming. The global Hailuo Video web app is one of the official interfaces through which users can access H3. However, the H3 API model is MiniMax-H3; it is not called MiniMax-Hailuo-3, and it does not use the older V1 Hailuo endpoint.

MiniMax describes the lineage this way: Hailuo 01 established the first-generation system; Hailuo 02 concentrated on architecture, data quality, efficiency, and scaling; H3 changes the design goal toward generalizing across tasks and modalities. The official pricing documentation now places Hailuo 2.3, Hailuo 2.3 Fast, and Hailuo 02 in its legacy-model table, while H3 uses the new Video Generation V2 API.

For detailed coverage of the previous family, see our independent pages for MiniMax Hailuo 2.3, Hailuo 2.3 Fast, and Hailuo 02.

MiniMax H3 vs Hailuo 2.3 and earlier models

QuestionMiniMax H3Hailuo 2.3Hailuo 2.3 Fast
Official status in current pricing docsCurrentLegacyLegacy
API generation endpoint/v2/video_generation/v1/video_generation/v1/video_generation
Documented input modesText; first/last frames; image, video, and audio referencesText-to-video and image-to-videoImage-to-video in the reviewed official catalog
Output resolution768P or 2K768P or 1080P768P or 1080P
Output duration4-15 seconds6 or 10 seconds, depending on resolution6 or 10 seconds, depending on resolution
Native synchronized audioYes, 32 kHz stereoNot documented as part of generationNot documented as part of generation
Public model weightsH3-Base weights released under a restricted community licenseNo official weights identifiedNo official weights identified
10-second 768P output price$0.80 before reference-input or Context-IR charges$0.56 per generated clip$0.32 per generated clip

H3 is not automatically the cheapest MiniMax option. At the published 10-second 768P rates, Hailuo 2.3 and Hailuo 2.3 Fast cost less. H3’s practical advantage is the broader multimodal workflow, native audio, longer duration, and hosted 2K output—not a universal lower price than its predecessors.

MiniMax H3 vs H3 Max: which current video model fits?

H3 is no longer MiniMax’s only current video model. MiniMax’s video guide now lists two: H3 and MiniMax H3 Max (MiniMax-H3-Max), which MiniMax describes as jointly released with fal.ai and post-trained by fal.ai on H3 for high-speed generation. Both run on the same Video Generation V2 endpoint, where MiniMax’s docs say “To use MiniMax H3 or MiniMax H3 Max, please select the Pay-as-you-go API.” You choose between them with the model field. On the subscription side, MiniMax’s M Plan lists the “H3 video model” in its Explore and Build tiers and does not mention H3 Max.

Decision pointMiniMax H3MiniMax H3 Max
Resolutions768P or 2K480P or 768P; 2K is not supported
Durations4–15 seconds5–15 seconds
Output price$0.08/s at 768P; $0.13/s at 2K$0.05/s at 480P; $0.08/s at 768P
Free reference images per request5, then $0.04 each2, then $0.074 each
Input video at 768P output$0.08/s$0.143/s
768P-to-2K regenerationDocumentedNot documented
Main reason to pick it2K masters, 4-second clips, cheaper reference-heavy requestsFaster turnaround and the cheapest V2 tier for drafts

In short, a plain 10-second clip at 768P costs $0.80 on either model, while H3 Max at 480P brings it down to $0.50. Once you add several reference images or a motion-reference video at 768P, H3 usually works out cheaper. Our dedicated MiniMax H3 Max guide covers its pricing, speed claims, free-access question, and worked cost examples in full.

MiniMax H3 input modes and limits

The official Video Generation V2 reference requires every request to contain a non-empty text item. Media can then be supplied as first/last frames or as reference material.

ModeDocumented content combination
Text-to-videoOne required text prompt
First-frame image-to-videoText plus one image with role=first_frame
Last-frame image-to-videoText plus one image with role=last_frame
First-and-last-frame videoText plus two images labeled as first and last frame
Reference-to-videoText plus any supported mix of reference images, reference videos, and reference audio

First/last-frame mode and reference mode are mutually exclusive in one request. If the content array contains a reference image, video, or audio role, it cannot also contain first-frame or last-frame roles.

Input ruleOfficial limit
Total request body64 MB maximum; MiniMax recommends public URLs instead of Base64 for large files
Reference imagesUp to 9; JPG, JPEG, PNG, WebP, HEIC, or HEIF
Single image size30 MB maximum
Image dimensionsWidth and height each between 256 and 5,760 pixels
Image aspect ratioWidth/height between 0.4 and 2.5
Reference videosUp to 3 MP4 or MOV clips; each 2-15 seconds, total duration no more than 15 seconds
Single video size50 MB maximum
Reference video codecsH.264/AVC or H.265/HEVC video; AAC or MP3 audio
Reference video frame rate23.976-60 FPS
Reference audioUp to 3 WAV or MP3 clips; each 2-15 seconds, total duration no more than 15 seconds
Single audio size15 MB maximum
Mixed reference files12 files in total across reference types, per the official Video Generation guide
Prompt lengthUp to 7,000 characters

For aspect ratios, text-to-video requires an explicit value and cannot use adaptive. Available documented values are 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Image-to-video follows the input image. Reference-to-video can use adaptive or a supported explicit ratio.

2K video and native stereo audio

What “native audio” means

H3 jointly predicts video and audio latents, then decodes them into video and 32 kHz stereo sound. This differs from a workflow that generates a silent clip first and attaches unrelated text-to-speech, stock music, or effects later. A prompt can describe dialogue, environmental sound, effects, and non-diegetic music as part of the intended scene.

Native generation does not guarantee perfect lip synchronization, pronunciation, factual speech, music rights, or clean mixing. Review the entire soundtrack and every frame before publication. MiniMax documents stable dialogue support for 11 languages, but “stable support” is a provider statement, not proof that every accent, name, or multilingual scene will be accurate.

How H3 produces 2K output

In the hosted API, users can request 2K directly. In the documented architecture, H3-Base first produces a 768P audiovisual result. H3-Regenerate-2K then combines that result with the original context and regenerates the video at higher resolution. MiniMax describes this as in-context regeneration rather than conventional super-resolution.

This distinction matters for local deployment: the released H3-Base weights cover 768P generation. The H3-Regenerate-2K module is not yet open, so a completely local open-weight path to the official 2K workflow is not currently documented.

How to access MiniMax H3

Access routeBest forImportant note
Hailuo Video web appCreators who want a graphical interfacePlans, credits, availability, and UI controls can differ by region and account
MiniMax Design desktopDesktop creative and agent workflowsThe former hub.minimax.io address now redirects to MiniMax Design. MiniMax’s H3 feature page advertised three free H3 generations for MiniMax Design desktop when checked on September 24, 2026; eligibility and promotions can change
MiniMax Code desktop v3.0.66 or laterDesktop workflows that combine H3 video generation with Skills, Browser Use, Remote Control, and GoalMiniMax documents H3 integration in Code. MiniMax Code is included with M Plan, whose Explore and Build tiers list the “H3 video model”; Go does not include a video model. Verify the live quota and account entitlement.
MiniMax Open Platform APIAutomated and production workflowsUses the V2 asynchronous API and pay-as-you-go billing
Hugging Face weightsLocal 768P research, evaluation, and deploymentRestricted community license; not the full hosted 2K system
Official GitHub repositoryDeployment instructions, workflows, examples, and prompt guidanceRead the license and hardware implications before downloading

For authentication, account setup, errors, and general production considerations, see our independent MiniMax API guide. For the difference between API billing and consumer plans, see our MiniMax pricing guide.

MiniMax H3 V2 API workflow

H3 generation is asynchronous:

  1. Send a request to POST /v2/video_generation.
  2. Store the returned task_id.
  3. Query GET /v2/query/video_generation/{task_id} or use a verified callback URL.
  4. When the status becomes succeeded, read the output URL from the task response.

The following is an original implementation example based on the official request schema. It is not a record of a paid test performed by this site:

curl --request POST \
  --url https://api.minimax.io/v2/video_generation \
  --header "Authorization: Bearer $MINIMAX_API_KEY" \
  --header "Content-Type: application/json" \
  --data '{
    "model": "MiniMax-H3",
    "content": [
      {
        "type": "text",
        "text": "A 10-second cinematic product film of a matte-black smartwatch resting on wet volcanic stone at dawn. The camera makes one slow clockwise orbit while droplets move across the glass. Keep the watch face and logo stable. Native audio: distant wind, soft water drops, and a restrained electronic pulse."
      }
    ],
    "resolution": "2K",
    "duration": 10,
    "ratio": "16:9"
  }'

Do not place an API key directly in client-side code or publish it in a repository. Use a server-side secret store, handle 4xx/5xx errors, enforce spending controls, and design for queues, retries, callbacks, moderation rejections, and output review.

When rechecked on September 24, 2026, the official rate-limit page listed MiniMax-H3 on Video Generation V2 at 300 requests per minute with a maximum of 30 in-flight tasks. The page no longer showed separate free-tier and paid-tier concurrency figures, and it did not yet list H3 Max. Account-specific limits can change, so handle 429 responses with backoff.

MiniMax H3 pricing

Pricing last verified: September 24, 2026. The official pay-as-you-go price list bills H3 output by generated second.

OutputOfficial list price
MiniMax H3 at 768P$0.08 per generated second
MiniMax H3 at 2K$0.13 per generated second
Regenerate an eligible 768P H3 video to 2K$0.05 per regenerated second, with applicable inputs billed again

Worked output-cost examples

Duration768P output2K output
5 seconds$0.40$0.65
10 seconds$0.80$1.30
15 seconds$1.20$1.95

These examples multiply duration by the published per-second output price. They exclude reference-material charges, H3-Context-IR token charges, taxes, currency conversion, failed retries outside the provider’s billing rules, consumer subscriptions, promotions, and third-party markups.

Reference-input and preprocessing charges

ItemPublished billing rule
Reference audioFree
Reference imagesFirst 5 free; $0.04 for each additional image
Reference videoBilled by input-video duration at $0.08/second for 768P output or $0.13/second for 2K output
H3-Context-IR$0.90 per million input tokens and $3.60 per million output tokens
768P-to-2K regeneration inputsThe original task’s input materials are billed again: audio free; first 5 images free, then $0.025 per image; input video $0.05 per second

Confirm the live billing page before a large batch. A video package or subscription that covers older Hailuo models should not be assumed to cover H3 unless the current checkout and plan terms say so explicitly.

MiniMax H3 open weights and local deployment

MiniMax released H3 weights on August 2, 2026 through the official MiniMaxAI/MiniMax-H3 repository. “Open weights” is the accurate short description; “fully open source” or “free for unrestricted commercial use worldwide” is not.

What is and is not included

H3 moduleRolePublic-weight status
H3-Context-IRInterprets and restructures complex multimodal contextHosted service; not included in the open-weight release
H3-BaseGenerates synchronized video and audio at 768PWeights released
H3-Regenerate-2KRegenerates the 768P result with the original context at 2KNot yet open; official API available
Sparse-attention implementationReduces long-sequence compute costNot included in the initial release; full-attention inference is available

Released H3-Base checkpoints

CheckpointSupported workflowPrecision
H3-Base-FL2VAText-to-audio-video; first frame, last frame, or first-and-last-frame to audio-videoBF16
H3-Base-Ref2VAText plus reference images, video, and/or audio to audio-videoBF16

The repository documents SGLang, vLLM, Diffusers, and ComfyUI workflows. Its example SGLang command uses four GPUs, but MiniMax does not state that this exact count or a specific GPU model is a universal minimum. H3-Omni-Transformer is described as a 33-billion-parameter dense model. Local deployment therefore requires serious hardware planning; do not purchase hardware based on a single example command.

MiniMax H3 Community License: important restrictions

The weights are governed by the MiniMax H3 Community License Agreement, not MIT, Apache 2.0, or another standard permissive open-source license.

  • Territory: the agreement excludes the United States, European Union, United Kingdom, and Republic of Korea. Deployment there requires a separate license from MiniMax.
  • Large commercial use: commercial products or services generating more than $20 million in annual revenue require prior written authorization.
  • Commercial attribution: a commercial product or service using H3 must prominently display “MiniMax H3” in its user interface.
  • Model improvement restriction: H3 works or outputs may not be used to improve another AI model, except H3 or its derivatives.
  • Use controls: licensees must follow the agreement, its Acceptable Use Policy, applicable law, and the required safeguards for hosted access.

This is a plain-language summary, not legal advice and not a substitute for reading the complete license. Organizations should obtain qualified legal review before local deployment, distribution, fine-tuning, or commercial integration.

MiniMax H3 examples

Official MiniMax examples

The following files are official provider examples hosted in MiniMax’s GitHub repository. They are useful for understanding the intended workflows, but they are curated vendor material—not independent evidence of average quality, latency, cost, or failure rate:

Original prompts to try

These prompts were written for this guide around H3’s documented controls. They have not yet been run as a controlled benchmark, so they should be treated as starting points rather than guaranteed recipes.

  • Product and native sound: “10 seconds, 16:9. A brushed-aluminum espresso machine on a dark stone counter at sunrise. Start with a macro shot of condensation, then make one slow pull-back as the machine pours into a white cup. Preserve the logo and button layout. Native audio: quiet room tone, switch click, steam hiss, liquid pouring, no voice and no music.”
  • First-and-last-frame transition: “Use Image 1 as the exact opening composition and Image 2 as the exact final composition. The empty studio gradually transforms into the furnished room through one continuous camera move. Objects should appear through physically plausible assembly rather than a dissolve. Keep wall dimensions, window position, and daylight direction consistent. Subtle construction sounds only.”
  • Multimodal motion reference: “Use the character identity and clothing from Image 1, the camera path and body timing from Video 2, and the beat structure from Audio 3. Place the performance in the neon train platform from Image 4. Preserve facial identity and outfit colors. Align major turns to the audio accents, with environmental station ambience mixed below the reference track.”

Limitations, review checks, and responsible use

Documented boundaries

  • 15-second maximum: H3 is still a short-clip generator. Longer work requires editing, continuation, or multiple shots.
  • Not a completely open 2K system: local H3-Base produces 768P; Context-IR and the 2K regeneration stage remain hosted or unreleased as weights.
  • Restricted license: public weights do not mean unrestricted global deployment or standard open-source licensing.
  • No independent H3 benchmark here yet: provider claims about quality, text rendering, brand rendering, and price-performance have not been independently reproduced by this site.
  • No fixed 2K dimensions published: use the returned task metadata instead of assuming a universal pixel size.
  • Input costs can matter: reference video, extra images, Context-IR, and regeneration can add to the output-only examples.
  • Interface differences: API capability does not guarantee that every feature appears in every Hailuo interface, plan, country, account, or third-party provider.

Review every output

Generative audiovisual output can contain identity drift, deformed anatomy, changing object counts, unstable text and logos, discontinuous motion, flicker, inaccurate dialogue, audio artifacts, unintended music, or unsafe background details. Review the full clip frame by frame and listen on headphones before publication. For product or brand work, compare packaging, colors, text, marks, and factual features against approved source material.

Rights, consent, privacy, and disclosure

  • Upload only media that you own, license, or are authorized to process.
  • Obtain appropriate consent before reproducing or animating an identifiable person, voice, performance, or private location.
  • Do not use H3 for impersonation, fraud, harassment, sexual exploitation, deceptive evidence, or unlawful surveillance.
  • Check copyright, trademark, publicity, music, advertising, and synthetic-media rules in the intended territory and platform.
  • Disclose AI-generated or materially altered media when required by law, platform policy, contract, or context.
  • Do not upload confidential, regulated, or sensitive material until your organization has approved MiniMax’s data flow, retention, account controls, and contract.

See our sensitive-data checklist and the provider’s current terms before using real customer, employee, patient, student, or unreleased commercial material.

MiniMax H3 FAQ

What is the exact MiniMax H3 API model ID?

Use MiniMax-H3 with the Video Generation V2 endpoint at POST https://api.minimax.io/v2/video_generation. Preserve the capitalization and hyphenation exactly.

Is MiniMax H3 free?

The API has published pay-as-you-go prices. MiniMax’s official H3 feature page advertised three free H3 generations through the MiniMax Design desktop app when this page was verified on September 24, 2026, but promotions, eligibility, regions, queues, and plan terms can change. Do not assume unlimited free use. For the separate question of free H3 Max access, see our H3 Max guide.

Is MiniMax H3 open source?

H3-Base weights and code have been publicly released, but “fully open source” is misleading. The weights use the restricted MiniMax H3 Community License, H3-Context-IR is hosted, H3-Regenerate-2K is not yet open, and the initial release does not include the sparse-attention implementation.

Can MiniMax H3 run locally?

Yes, H3-Base can run locally for 768P generation using the released FL2VA or Ref2VA checkpoints. MiniMax documents SGLang, vLLM, Diffusers, and ComfyUI routes. Read the license first and plan for a large 33B dense audiovisual model; the official example configuration is not a universal hardware guarantee.

Do the open weights generate 2K locally?

Not through a completely open official stack at the time of verification. H3-Base generates 768P. The official 2K workflow uses H3-Regenerate-2K, which is available through MiniMax’s hosted API but has not yet been released as open weights.

Does MiniMax H3 generate audio with the video?

Yes. H3 jointly generates video and native 32 kHz stereo audio. The sound can include speech, environmental sound, effects, and music described by the prompt, subject to model performance and provider controls.

Is H3 the same as Hailuo 3?

No official API model called “Hailuo 3” was identified in the sources reviewed. MiniMax-H3 is the API model ID, while Hailuo Video is an official application that provides access to it. Earlier models use IDs such as MiniMax-Hailuo-2.3.

Is H3 better than Hailuo 2.3?

H3 has objectively broader documented inputs, native stereo audio, 4-15 second duration, hosted 2K output, and released base weights. Hailuo 2.3 is cheaper in the published 10-second 768P comparison. “Better” for visual quality, latency, identity consistency, text, and prompt following requires a controlled benchmark using the same briefs and review criteria.

What is the difference between MiniMax H3 and H3 Max?

H3 Max is a speed-optimized model that fal.ai post-trained on H3. It outputs 480P or 768P for 5–15 seconds at $0.05 or $0.08 per second, while H3 outputs 768P or 2K for 4–15 seconds at $0.08 or $0.13 per second. H3 also includes more free reference images and documents 2K regeneration. See the H3 Max guide for details.

Should I use H3 or H3 Max?

Use H3 when you need 2K, 4-second clips, or many reference inputs at 768P. Use H3 Max for fast iteration and low-cost 480P drafts. For plain 768P text-to-video, both cost $0.08 per second, so the choice comes down to speed and the look you prefer.

Official sources and verification notes

Prices, limits, model availability, promotions, licenses, and interfaces can change without notice. Corrections can be submitted through our contact and corrections page. Our broader sourcing and testing approach is described on About & Methodology.

Update log

DateVerified change
July 31, 2026MiniMax announced H3 and its hosted 2K, native-audio, multimodal generation capabilities.
August 2, 2026MiniMax H3 Community License date; H3-Base weights released through the official Hugging Face/GitHub project.
August 11, 2026This independent guide was created and checked against the current announcement, API, pricing, rate-limit, repository, and license sources.
August 20, 2026Added MiniMax Code v3.0.66, released August 19, as a documented H3 desktop access route; kept that product integration separate from Token Plan coverage, whose current table explicitly excludes H3.
September 24, 2026Added the H3 vs H3 Max comparison and FAQ entries; updated rate limits (300 RPM, 30 in-flight tasks), the MiniMax Design access route that replaced hub.minimax.io, the 12-file cap and 7,000-character prompt limit from the official guide, and regeneration input prices. H3 output and input prices were unchanged.
September 30, 2026Replaced the Token Plan exclusion notes with MiniMax’s new M Plan wording: Explore and Build list the “H3 video model,” Go has no video model, and the V2 API docs still say to select the Pay-as-you-go API for H3 and H3 Max.

Conclusion: MiniMax H3 is the MiniMax choice when a workflow needs one model for short video, native stereo audio, multimodal references, and hosted 2K output; for faster, lower-cost drafts at up to 768P, compare it with H3 Max. It is a larger change than Hailuo 2.3, but it also introduces a new API, per-second pricing, substantial local hardware requirements, and a restrictive open-weight license. Choose the hosted route for the complete managed workflow, or local H3-Base for controlled 768P evaluation after reviewing the license, hardware, privacy, and safety requirements.

More on Hailuo: see the Hailuo AI hub for every model compared side by side, or Is Hailuo AI free? for what each access route actually costs.