Last verified: September 30, 2026 (subscription prices and the M3.1-Flash-Preview comparison). Other M3 details were last fact-checked August 26, 2026.
MiniMax M3 is an open-weight, multimodal Mixture-of-Experts model built for coding, tool use, long-running agents, and large-context workloads. Released on June 1, 2026, it has approximately 428 billion total parameters, activates about 23 billion parameters per token, accepts text, images, and video, and offers an API context window of up to one million tokens. Its weights are also available for self-hosting. (Official MiniMax M3 page; Hugging Face model card)
The model’s strongest practical proposition is not any single benchmark score. It is the combination of long context, native multimodal input, agent-oriented reasoning, downloadable weights, OpenAI- and Anthropic-compatible APIs, and relatively low first-party API prices.
There are important qualifications, however. Many headline benchmark results come from MiniMax’s own evaluations, the downloadable weights use a commercially restricted custom license rather than a standard permissive open-source license, and full-scale self-hosting remains a data-center-class undertaking.
Bottom line: MiniMax M3 is a serious option for coding agents, repository-scale analysis, long-running tool workflows, and multimodal document work. It should be tested against real workloads rather than selected solely from the advertised one-million-token limit or vendor benchmark charts.
MiniMax M3 at a Glance
| Specification | Verified detail |
|---|---|
| Release date | June 1, 2026 |
| API model ID | MiniMax-M3 |
| Architecture | Native multimodal Mixture of Experts with MiniMax Sparse Attention |
| Total parameters | Approximately 428 billion |
| Active parameters | Approximately 23 billion per token |
| Input modalities | Text, images, and video |
| Output modality | Text |
| API context window | Up to 1,000,000 tokens; the product page states a guaranteed minimum of 512,000 |
| Maximum output | 524,288 tokens (hard maximum); MiniMax recommends 131,072 tokens for M3 |
| Reasoning modes | Adaptive thinking or thinking disabled |
| API formats | Anthropic-compatible and OpenAI-compatible |
| Weights | Available on Hugging Face; official GitHub repositories provide deployment code and supporting resources. |
| Model license | MiniMax Community License |
| Standard API price at up to 512K input | $0.30 per million input tokens and $1.20 per million output tokens |
| Standard API price above 512K input | $0.60 per million input tokens and $2.40 per million output tokens |
The one-million-token limit covers the combined input and output budget rather than allowing one million input tokens plus an additional maximum-length answer. The API documentation also distinguishes the overall 1M ceiling from the separately configurable output limit.
What Is MiniMax M3?
MiniMax M3 is the latest M-series model listed in MiniMax’s current API documentation. It is positioned primarily for software engineering, tool-driven agents, structured task execution, long documents, large codebases, and multimodal understanding. Unlike the earlier text-focused M-series releases, M3 can process images and video alongside text. (MiniMax model invocation documentation)
The model is a Mixture of Experts, or MoE. Instead of activating every parameter for every token, a routing mechanism selects a smaller group of experts. MiniMax reports approximately 428 billion parameters in total but about 23 billion active for each token. This can reduce inference compute compared with a similarly sized dense model.
The “23B active” figure does not make MiniMax M3 equivalent to an ordinary 23-billion-parameter model for deployment. The complete expert weights still have to be stored and served. Sparse activation lowers the amount of computation performed for each token, but it does not make the other hundreds of billions of parameters disappear from storage or memory planning. That distinction is one reason full-weight deployments use multi-GPU systems despite the relatively modest active-parameter figure.
Native multimodality
MiniMax says M3 was trained with mixed modalities from the beginning rather than having vision support attached after text training. Through the current APIs, it can accept:
- Text
- JPEG, PNG, GIF, and WEBP images
- MP4, AVI, MOV, and MKV video
- Tool definitions and tool results
- Multi-turn reasoning content
The output is text, not generated images or video. The Anthropic-compatible API currently permits direct images of up to 10 MB and direct videos of up to 50 MB, while larger videos can be passed through the Files API under separate limits. These limits should be checked again before building a production ingestion pipeline.
Coding, agents, and computer use
M3 is designed to work inside an agent harness that can read files, run commands, invoke tools, inspect results, and continue reasoning between actions. MiniMax calls this interleaved thinking: the model can reassess its state after each tool result rather than planning once and executing a fixed sequence blindly.
MiniMax also presents M3 as a computer-use model through MiniMax Code. This should not be confused with an ordinary API request gaining unrestricted desktop access. Computer use still requires an orchestration layer that can capture screens, expose actions, manage permissions, execute clicks or keystrokes, and enforce security boundaries.
How MiniMax Sparse Attention Works
A conventional full-attention mechanism compares tokens across the entire sequence. As the sequence grows, the work required by full attention increases quadratically, making very long contexts costly.
MiniMax M3 replaces full attention with MiniMax Sparse Attention, or MSA. According to the MSA technical paper, it is a blockwise sparse-attention system built on Grouped Query Attention:
- A lightweight index branch scores blocks of keys and values.
- It selects a limited set of relevant blocks for each GQA group.
- The main branch performs exact attention only over the selected blocks.
- GPU kernels arrange the work to improve memory access and hardware utilization.
The intent is to preserve access to relevant information without repeatedly attending to every previous token.
What the reported efficiency figures mean
MiniMax’s model card reports that M3 achieves nine-times faster prefill and 15-times faster decoding than M2 at a one-million-token context, while reducing per-token compute to one-twentieth of the previous model’s level. These are MiniMax-reported comparisons against M2, not universal guarantees for every server, provider, prompt, or application.
The separate MSA paper reports a different controlled experiment: on a 109-billion-parameter model running on H800 hardware, MSA reduced attention compute by 28.4 times at one million tokens and delivered 14.2-times faster prefill and 7.6-times faster decoding than its GQA comparison. The figures differ because the baselines, model sizes, and test conditions differ.
This distinction matters. The paper’s primary quality and efficiency experiments were performed on a 109B experimental MoE model, not the complete 428B MiniMax M3 checkpoint. The experiments support the MSA design, but they are not an independent end-to-end validation of every production-model claim.
MSA does not eliminate all long-context costs
Sparse attention addresses a major source of compute, but a one-million-token application still has to manage:
- Input transfer and tokenization
- Vision and video preprocessing
- Key-value cache memory
- Prompt ingestion latency
- Model output and reasoning tokens
- Tool-call history
- Retrieval quality across distant evidence
- Provider-specific serving limits
The MSA kernels are separately available through an MIT-licensed repository, but that permissive kernel license does not replace the custom license applied to the full M3 model weights. (Official MSA repository)
The 1M Context Window: Capability and Caveats
MiniMax’s API documentation lists a 1,000,000-token context window. Its product page describes support for “up to 1M” tokens while also stating a guaranteed minimum of 512K. The maximum token count is the combined total of input and output tokens.
For example, a request containing 900,000 input tokens cannot also produce a 512,000-token answer. The requested output must fit inside the remaining context budget.
The current generation limits are:
- Recommended maximum output for M3: 131,072 tokens
- Hard maximum output setting: 524,288 tokens
- Total context: Input and output together must fit within 1,000,000 tokens
A context ceiling is not a quality guarantee
An API accepting one million tokens does not prove that every detail remains equally retrievable across that entire sequence. Long-context performance depends on the type of information, where evidence appears, how much irrelevant material surrounds it, the prompt structure, and the reasoning task.
MiniMax’s own prompting guidance recommends indexing and clearly delimiting long source packages, putting the task after the source material, and asking the model to identify or summarize relevant evidence before producing its final answer.
For repository or document analysis, a better workflow is usually:
- Remove generated files, binaries, duplicate documents, and irrelevant logs.
- Give each source a clear name, date, and boundary.
- Add a repository map or document index.
- Put the actual question after the source material.
- Require file, section, or quotation references.
- Use the token-count endpoint before sending the request.
- Test retrieval using facts located near the beginning, middle, and end of the context.
- Compare full-context performance with retrieval-augmented or staged processing.
The ability to fit an entire repository into one request can be useful, but sending everything indiscriminately may increase cost and noise without improving accuracy.
Crossing 512K changes the price
MiniMax’s first-party pricing switches to a higher tier when input tokens exceed 512K. Cache-hit tokens count toward that input threshold. A 513K-token input can therefore cost more per input and output token than a 500K-token input, even when the generated answer is short.
This creates a practical optimization boundary. Before crossing it, consider whether indexing, deduplication, retrieval, or staged summarization can reduce the request without losing necessary evidence.
MiniMax M3 Benchmarks: Vendor Claims vs. Independent Results
MiniMax reports strong performance across coding and agentic benchmarks. The most frequently cited figures include 59.0% on SWE-Bench Pro, 66.0% on Terminal-Bench 2.1, 34.8% on SWE-fficiency, 28.8% on KernelBench Hard, and 74.2% on MCP Atlas.
These are useful signals, but they should not be presented as independent measurements. MiniMax’s launch report explains that several evaluations were run on internal infrastructure, often with Claude Code or another agent scaffold, custom timeouts, specific sandboxes, modified system prompts, or model-based judges.
| Evaluation | Result | Who ran it? | Important qualification |
|---|---|---|---|
| SWE-Bench Pro | 59.0% | MiniMax | Internal infrastructure using Claude Code as the scaffold; MiniMax says its logic was aligned with the official evaluation |
| Terminal-Bench 2.1 | 66.0% | MiniMax | Internal 8-core, 16 GB sandbox with a two-hour timeout and Terminus 2; some competitor results came from the official leaderboard while others were rerun internally |
| SWE-fficiency | 34.8% | MiniMax | Internal testing with the open dataset and Claude Code scaffold |
| KernelBench Hard | 28.8% | MiniMax | Internal evaluation on Blackwell hardware; the score depends on throughput relative to theoretical hardware peak |
| MCP Atlas | 74.2% | MiniMax | Official codebase, but the public-set evaluation used Gemini 2.5 Pro as a scoring model |
| Artificial Analysis Intelligence Index v4.1 | 44 | Artificial Analysis | Independently run composite covering nine evaluations, not a dedicated coding-only score |
| Output speed | 96.8 tokens/second | Artificial Analysis | Measured through MiniMax’s first-party API; actual speed varies by prompt, region, load, and service tier |
| Time to first token | 1.71 seconds | Artificial Analysis | Provider-specific API measurement rather than a self-hosting result |
The complete methodology notes are available in the MiniMax launch report, while the current independent figures are listed on Artificial Analysis.
Why some articles show an Artificial Analysis score of 55
Early June 2026 coverage reported an Artificial Analysis Intelligence Index score of 55 for M3. The current page shows 44 under Intelligence Index v4.1.
That does not necessarily mean the model became worse. Artificial Analysis introduced v4.1 on June 15, changing and reweighting the underlying evaluation suite toward newer agentic workloads. Scores from different index versions should not be plotted as if they came from the same test.
Under the current v4.1 ranking, Artificial Analysis places GLM-5.2 at 51, MiniMax M3 at 44, and DeepSeek V4 Pro at 44 among the leading open-weight models.
Long-running demonstrations are promising, but first-party
MiniMax also describes several extended autonomous tasks:
- A nearly 12-hour reproduction of experiments from an ICLR paper
- A roughly 24-hour CUDA optimization run involving 147 benchmark submissions and 1,959 tool calls
- A reported 9.4-times speed improvement in that CUDA task
- A 12-hour post-training experiment in which the agent selected data, trained models, and evaluated results
These demonstrations are relevant because long-running agents can fail through context drift even when they perform well on single-turn coding questions. However, they were organized and reported by MiniMax. Treat them as first-party case studies rather than independently reproduced evidence.
How to evaluate M3 for your workload
A practical evaluation should test the complete system, not just the raw model:
- Use the same agent harness, tools, permissions, and timeout for every model.
- Separate time to first token, reasoning time, output speed, and total task time.
- Record total input, cache-read, reasoning, and answer tokens.
- Require tests to pass rather than judging code by appearance.
- Run each task more than once.
- Include regression repair, repository navigation, ambiguous requirements, and tool failures.
- Score whether the model knows when to stop.
- Measure human correction time.
- Test both thinking-enabled and thinking-disabled modes.
- Compare cost per successful task, not only cost per million tokens.
A model that is slightly more expensive per token can be cheaper per completed task if it needs fewer retries. The reverse can also be true when verbose reasoning consumes large output budgets.
MiniMax M3 API Pricing
The following prices were displayed in MiniMax’s first-party pay-as-you-go pricing documentation on July 10, 2026.
| Service tier | Input length | Input | Output | Prompt-cache read |
|---|---|---|---|---|
| Standard | Up to 512K input tokens | $0.30/M | $1.20/M | $0.06/M |
| Standard | Above 512K input tokens | $0.60/M | $2.40/M | $0.12/M |
| Priority | Up to 512K input tokens | $0.45/M | $1.80/M | $0.09/M |
| Priority | Above 512K input tokens | $0.90/M | $3.60/M | $0.18/M |
M means one million tokens. Priority service is priced at 1.5 times the standard rate and is intended to provide preferential request admission, faster responses, and improved reliability under load. The pricing page currently labels the displayed rates as a permanent 50% discount from higher struck-through list prices.
Two cost examples
Assume standard service without cache hits.
Example 1: A repository-analysis request with 200K input tokens and 20K output tokens
Input: 0.20 × $0.30 = $0.060
Output: 0.02 × $1.20 = $0.024
Total: $0.084
Example 2: A long-context request with 700K input tokens and 30K output tokens
Input: 0.70 × $0.60 = $0.420
Output: 0.03 × $2.40 = $0.072
Total: $0.492
The second request is charged entirely under the above-512K tier because its input crosses the threshold.
These are hypothetical calculations, not observed invoices. Actual cost can include cache behavior, image or video tokens, reasoning output, retries, tool-call history, and third-party provider markups.
Prompt caching
Prompt caching can materially reduce the cost of repeated system prompts, repository snapshots, specifications, or policy documents. Cache-read tokens are billed below ordinary input rates. MiniMax’s documentation says passive prompt caching does not add a separate cache-write fee, while explicit Anthropic-style caching has different behavior.
Caching is most useful when a large, stable prefix is reused across multiple requests. It provides less value when each request changes the beginning of the prompt or sends a completely different repository state.
Token Plan subscriptions
MiniMax offers Token Plans for interactive developer workflows. Its Token Plan price page lists Plus, Max, and Ultra at $22, $55, and $132 per month. On September 30, 2026, MiniMax introduced M Plan (Go, Explore, Build) at the same prices. The final checkout amount controls.
The same live purchase data shows approximate M3 monthly-use references of about 1.7B, 5.1B, and 12.5B tokens. These are M3-specific estimates, not guaranteed universal balances. Actual entitlement is governed by the shared usage bar, a five-hour rolling window, a weekly limit, and the resource consumed.
Current Token Plan coverage lists text, image, and speech resources; Music is not listed. MiniMax-H3, Voice Design, and Rapid Voice Cloning are excluded. Subscription Keys and ordinary pay-as-you-go API keys are not interchangeable, and MiniMax recommends pay-as-you-go for production use.
M3 vs M3.1-Flash-Preview
On September 27, 2026, MiniMax released MiniMax-M3.1-Flash-Preview, a newer M-series model with the same 1,000,000-token context and text, image, and video input. MiniMax’s model table now lists it first and says M3 “also remains available.” The practical differences are about access and control, not published performance:
| Item | MiniMax M3 | M3.1-Flash-Preview |
|---|---|---|
| Access | Pay-as-you-go API, subscription plans, MiniMax Code, open weights | Subscription plan (Token Plan / M Plan) and MiniMax Code only, for now |
| Pay-as-you-go price | $0.30 input / $1.20 output per 1M tokens (Standard, up to 512K input) | None published |
| Thinking | Can be enabled or disabled; defaults differ by endpoint | Always on; effort of low, medium, high, xhigh, or max (default max) |
| Benchmarks and speed | MiniMax-published results; about 100+ tps listed | None published yet |
For production apps that need metered billing, M3 remains the model with a published per-token price. The preview is for subscribers who want to try MiniMax’s newest coding model in MiniMax Code or their own tools.
How to Access MiniMax M3 Through the API
MiniMax supports both major compatibility formats:
| Format | Base URL | Best fit |
|---|---|---|
| Anthropic-compatible | https://api.minimax.io/anthropic | Recommended by MiniMax; supports native thinking blocks and interleaved tool reasoning |
| OpenAI-compatible | https://api.minimax.io/v1 | Easier migration from existing OpenAI SDK applications |
The model ID is MiniMax-M3.
Minimal Python example with the Anthropic SDK
import os
import anthropic
api_key = os.getenv("MINIMAX_API_KEY")
if not api_key:
raise RuntimeError("Set the MINIMAX_API_KEY environment variable.")
client = anthropic.Anthropic(
base_url="https://api.minimax.io/anthropic",
api_key=api_key,
)
message = client.messages.create(
model="MiniMax-M3",
max_tokens=4096,
thinking={"type": "adaptive"},
messages=[
{
"role": "user",
"content": (
"Review this Python function for correctness, security, "
"and maintainability. Return the blocking issues first."
),
}
],
)
answer_parts: list[str] = []
for block in message.content:
if block.type == "text":
answer_parts.append(block.text)
print("\n".join(answer_parts))
This follows MiniMax’s current Anthropic-compatible configuration and explicitly enables adaptive thinking rather than relying on an endpoint default. (Official model invocation guide)
Set thinking behavior explicitly
The current compatibility layers do not share the same default:
- In the Anthropic-compatible Messages API, omitting
thinkingdisables M3 thinking by default. - In the OpenAI-compatible Chat Completions API, omitting
thinkingenables adaptive thinking by default. - In the OpenAI-compatible Responses API, omitting
reasoningdisables reasoning output by default.
This can produce confusing latency, output, and cost differences during a migration. Production code should set reasoning behavior explicitly and test the precise endpoint being deployed.
Use adaptive thinking for debugging, planning, difficult coding, and multi-step tool work. Disable it for extraction, simple classification, low-latency completion, or other tasks that do not benefit from extended reasoning.
Preserve reasoning state during tool calls
For multi-step function calling, MiniMax instructs developers to append the model’s complete assistant response to the conversation history, including tool calls and thinking or reasoning_details fields. Dropping this state can degrade continuity or cause API-format errors in later turns.
A safe agent loop should therefore:
- Send the task and tool definitions.
- Store the complete assistant response.
- Validate tool arguments.
- Execute the tool within a restricted environment.
- Append the tool result.
- Send the full history back to M3.
- Continue until the model returns no further tool calls or a configured step limit is reached.
Always enforce application-side limits for total steps, wall-clock time, spending, file access, network access, and destructive actions. A model deciding to call a tool is not equivalent to receiving authorization to execute it.
Where MiniMax M3 Is Most Useful
Repository-scale coding work
The long context can hold large portions of a codebase, architectural documentation, recent diffs, test output, and issue history in one working session. Good candidate tasks include:
- Cross-file refactoring
- Migration planning
- Regression diagnosis
- Repository question answering
- Test generation
- Dependency analysis
- Long-running implementation tasks
The model still needs a competent harness. File search, code execution, test running, version-control inspection, and rollback procedures often matter as much as the underlying model.
Long-document analysis
M3 can process extensive technical standards, financial reports, policy collections, contracts, support archives, or research material. Its multimodal input also makes it relevant when documents contain diagrams, screenshots, charts, or scanned visual elements.
For high-stakes legal, financial, or medical material, long context does not remove the need for source citations, specialist review, and deterministic checks.
Visual software and interface work
Image input can help the model compare an implementation with a mockup, inspect screenshots, identify layout differences, or reason about visual error states. Video input may be useful for reviewing workflows, interface recordings, or long demonstrations.
Visual support should be benchmarked separately from text coding. A strong repository score does not automatically predict accurate visual grounding.
Tool-driven research and operations
Interleaved thinking makes M3 a candidate for agents that search, retrieve documents, call APIs, run calculations, update drafts, and verify intermediate results. The strongest workflows make tool boundaries explicit and require evidence before the model makes time-sensitive claims.
Can MiniMax M3 Be Self-Hosted?
Yes. The weights are available on the official Hugging Face repository, and MiniMax lists vLLM, SGLang, Transformers, KTransformers, and Unsloth as supported or recommended deployment paths. Quantized community and vendor variants are also available.
However, “available for local deployment” should not be interpreted as “easy to run on a normal workstation.”
The base checkpoint contains roughly 427–428 billion parameters. One published Lambda deployment example uses an NVIDIA HGX B200 system with eight B200 GPUs and tensor parallelism across all eight devices. This reflects Lambda’s tested deployment configuration rather than an official MiniMax minimum hardware requirement. Hardware requirements vary depending on quantization, runtime, context length, and deployment goals.
That is not necessarily a universal hardware minimum—quantization and alternative runtimes can change the requirements—but it is a useful indication of the model’s scale.
Why deployment requirements vary
Memory and performance depend on:
- Weight precision
- Quantization format
- Vision components
- Tensor or expert parallelism
- Inference framework
- MSA kernel support
- Batch size and concurrency
- Maximum sequence length
- Key-value cache precision
- Prompt-caching configuration
- Desired throughput
- Safety margin for production traffic
A quantized model may fit on significantly less hardware than the full-precision checkpoint, but it may also introduce quality loss, unsupported multimodal components, reduced context, slower offloading, or custom-kernel requirements.
Self-hosting checklist
Before choosing self-hosting, verify:
- The exact checkpoint and quantization license
- Full model and tokenizer compatibility
- Image and video support in the selected runtime
- Reasoning and tool-call parsers
- MSA support on the target GPU architecture
- Maximum tested context under realistic concurrency
- Weight and key-value cache memory
- Time to first token and sustained throughput
- Failure recovery and observability
- Security patching
- The cost of engineering and idle infrastructure
Self-hosting can improve control over data and deployment, but the first-party API will usually be simpler and cheaper for early evaluation or irregular traffic.
Is MiniMax M3 Open Source?
The most accurate description is open-weight with a custom, commercially restricted license.
The weights and model files can be downloaded, inspected, modified, and self-hosted. However, they are distributed under the MiniMax Community License, not a conventional permissive license such as Apache 2.0 or MIT.
The current license grants non-commercial permissions subject to retaining the copyright and permission notice. For commercial use, it requires users to:
- Display “Built with MiniMax M3” prominently on a relevant website, interface, blog post, about page, or product documentation.
- Obtain prior written authorization if the relevant products or services generate more than $20 million in yearly revenue.
- Send MiniMax a one-time notice when commercial revenue does not exceed that threshold.
The license defines commercial use broadly enough to include paid products, commercial APIs, hosted services, and commercially deployed fine-tuned or otherwise modified derivatives. It also contains a prohibited-use appendix.
These terms mean that:
- Open-weight is accurate.
- Self-hostable is accurate.
- Free for every commercial use without conditions is inaccurate.
- Permissively open source is inaccurate.
This is a summary, not legal advice. Organizations should review the current license, keep a dated copy of the version they evaluated, and obtain legal review before commercial deployment. API access may also be governed by separate platform terms.
Privacy and Production Due Diligence
A hosted M3 request sends prompts, source material, images, videos, code, and tool history to the selected API provider. Privacy conditions can differ between MiniMax’s first-party API and third-party hosts.
MiniMax’s public API privacy policy states that personal data may be retained for as long as necessary or permitted by law or to fulfil relevant purposes. The public policy does not provide one simple, universal zero-data-retention guarantee for every M3 request and customer configuration. (MiniMax API Privacy Policy)
Before processing confidential repositories or regulated data, obtain written answers covering:
- Whether prompts and outputs are used for model training
- Default and configurable retention periods
- Zero-data-retention eligibility
- Processing and storage regions
- Subprocessors
- Encryption
- Access logging
- Deletion procedures
- Incident notification
- Data-processing agreements
- Enterprise isolation
- Support personnel access
Do not transfer a third-party host’s privacy promise to MiniMax’s first-party API, or vice versa. Provider selection changes the operational privacy model even when the underlying weights are identical.
Self-hosting can keep inference inside an organization’s environment, but only when telemetry, logs, object storage, backups, monitoring systems, and external tools are configured consistently with that objective.
MiniMax M3 vs. alternatives: a dated evaluation framework
No model is the best choice for every workload, and a third-party leaderboard snapshot should not be treated as a permanent ranking. Recheck each provider’s current model catalog, prices, license, and the exact benchmark version on the day of evaluation.
| Decision question | Why M3 may fit | What to test against |
|---|---|---|
| Do you need one million tokens plus image and video input? | M3 documents all three and returns text | The current hosted multimodal flagship from each shortlisted provider |
| Do you need downloadable weights? | M3 weights are available under the MiniMax Community License | Exact competing checkpoint, model card, hardware fit, and license—not the provider’s hosted-brand claims |
| Is lowest raw token price the priority? | M3 has low posted rates at or below 512K input | Total cost per accepted result, including cache, long-context tier, tools, retries, latency, and review |
| Is coding or agent reliability the priority? | MiniMax targets coding, tool use, and long-running agents | Your repository tasks, tool schema, failure recovery, and time-to-merge rather than one headline benchmark |
| Do you need enterprise governance? | Hosted and self-managed routes offer different control | Current privacy terms, region, retention, access controls, support, and incident requirements |
For dated provider-specific checks, use this site’s current comparisons with ChatGPT, Gemini, Qwen, DeepSeek, Mistral, and Grok, then verify every volatile figure against the linked official source before procurement.
Strengths and Limitations
| Strengths | Limitations |
|---|---|
| One-million-token API ceiling | Product page qualifies it with a guaranteed minimum of 512K |
| Native text, image, and video input | Text output only |
| Strong focus on coding and long-horizon agents | Many headline benchmarks are vendor-run |
| Open weights and multiple deployment frameworks | Custom commercial-use restrictions |
| Approximately 23B active parameters per token | Complete 428B-class weights remain large to store and serve |
| Competitive first-party API pricing | Input above 512K doubles standard token rates |
| OpenAI- and Anthropic-compatible endpoints | Thinking defaults differ between endpoint formats |
| Adaptive and disabled thinking modes | Reasoning can increase latency and output-token cost |
| Prompt caching | Production behavior still requires workload-specific testing |
| Potential control through self-hosting | Practical full-scale hosting requires substantial hardware and expertise |
Who Should Consider MiniMax M3?
M3 is a strong candidate for teams that need several of its capabilities together:
- Coding agents that work across large repositories
- Long-running tool-use workflows
- Multimodal analysis of documents, screenshots, and video
- Large-context research or compliance review
- Open-weight deployment with organization-controlled infrastructure
- API experimentation where low token prices matter
- Teams already using OpenAI- or Anthropic-style SDKs
It is less suitable as a default choice when:
- The model must run on a small consumer machine.
- A standard permissive license is mandatory.
- The task only needs a small, fast model.
- The application cannot send data to a hosted provider and self-hosting is impractical.
- The organization needs mature enterprise guarantees before testing.
- The decision is being made solely from one vendor benchmark.
- A one-million-token prompt is being used to avoid proper retrieval and context management.
Frequently Asked Questions
Is MiniMax M3 open source?
Its weights are publicly downloadable, so “open-weight” is accurate. The model is not distributed under an unrestricted permissive open-source license. Its MiniMax Community License includes attribution, notification, revenue, and prior-authorization conditions for commercial use.
Is MiniMax M3 free?
Downloading the weights does not require paying an API token fee, but running them requires substantial compute. MiniMax’s hosted API is paid, and commercial self-hosting remains subject to the model license.
Does MiniMax M3 really support one million tokens?
The API documentation lists a total context window of 1,000,000 tokens. MiniMax’s model page describes support up to 1M and a guaranteed minimum of 512K. Input and output share the total limit, and inputs above 512K enter a higher pricing tier.
How much does the MiniMax M3 API cost?
At the standard service tier, current first-party prices are $0.30 per million input tokens and $1.20 per million output tokens for inputs up to 512K. Above 512K input tokens, the prices rise to $0.60 and $2.40. Priority service costs 1.5 times the standard rates. Prices were checked on July 10, 2026.
Can MiniMax M3 run locally?
Yes, but the full model is not a conventional laptop-sized LLM. The base checkpoint has about 428B total parameters, and one documented full deployment uses an eight-GPU NVIDIA HGX B200 system. Quantized deployments can reduce the requirement, but usable context, speed, multimodality, and accuracy depend on the exact runtime and checkpoint.
Does MiniMax M3 support images and video?
Yes. The current APIs accept text, image, and video input and produce text output. File formats and size limits depend on the API format and upload method.
Is MiniMax M3 compatible with the OpenAI API?
MiniMax provides an OpenAI-compatible endpoint at https://api.minimax.io/v1. It also provides an Anthropic-compatible endpoint, which MiniMax recommends for native thinking blocks and interleaved tool reasoning.
Is MiniMax M3 good for coding?
The available evidence indicates that coding and agentic work are core strengths. MiniMax reports competitive results on SWE-Bench Pro, Terminal-Bench 2.1, SWE-fficiency, and related evaluations, while Artificial Analysis places M3 among the leading current open-weight models. The exact answer still depends on the programming language, repository, harness, tools, and task duration.
Final Verdict
MiniMax M3 is one of the more capable open-weight multimodal models released in 2026, combining long context, multimodal input, agent-oriented capabilities, downloadable weights, and competitive API pricing.
Its most important advantage is the combination of these features. Teams do not have to choose separately between an open-weight coding model, a long-context model, and a multimodal input model.
Its most important limitation is that the headline proposition requires qualification. The 1M window is a maximum rather than proof of perfect one-million-token recall; many leading benchmark results are first-party; the weights carry commercial restrictions; and self-hosting a 428B-class model is operationally demanding.
MiniMax M3 deserves a controlled evaluation when the workload involves large repositories, long-lived agents, visual material, or substantial context. Test it with the same tools, timeouts, prompts, and success criteria used for competing models. Measure total cost per successful task, review the current license, and verify provider privacy terms before production deployment.
For teams that need a permissive license, minimal hardware, the strongest current independent composite score, or mature contractual guarantees, another model may be a better default. For teams specifically seeking an affordable, multimodal, long-context, open-weight foundation for coding and agent workflows, MiniMax M3 belongs near the top of the shortlist.
