MiniMax M2.7 is a text-and-tool-use model released by MiniMax on March 18, 2026 for software engineering, agent workflows, and professional document tasks. It remains listed in MiniMax’s main language-model catalog, but MiniMax-M3 has been the frontier M-series model since June 1, 2026.
Verified model status
Last verified: July 17, 2026. MiniMax’s official catalog lists MiniMax-M3, MiniMax-M2.7, and MiniMax-M2.7-highspeed outside the Legacy Models section. It places MiniMax-M2.5, M2.5-highspeed, M2.1, M2.1-highspeed, and M2 under Legacy Models. M2.7 is therefore still a documented, non-legacy option, but it should not be presented as MiniMax’s frontier M-series model.
This is an independent guide and is not affiliated with or endorsed by MiniMax. Product status, account access, prices, limits, and license terms can change; verify them with the linked MiniMax sources before deployment.
MiniMax M2.7 at a glance
| Official model ID | MiniMax-M2.7 |
|---|---|
| Faster hosted variant | MiniMax-M2.7-highspeed |
| Release date | March 18, 2026 |
| Catalog status verified July 17, 2026 | Listed outside the Legacy Models section; M3 is the frontier M-series model |
| Context limit | 204,800 tokens combined across input and output |
| Compatible input types | Text and tool-call content; no image or video input for M2.x through the compatible interfaces |
| Documented hosted speed | Approximately 60 tokens/second for M2.7 and 100 tokens/second for highspeed; these are documentation figures, not a latency guarantee |
| Hosted access | MiniMax Open Platform API; MiniMax also links M2.7 from its Agent and Token Plan materials |
| Local access | Public weights from MiniMax’s official repository and Hugging Face organization |
| Weight license | Non-commercial license with MIT-style permissions for permitted non-commercial uses; commercial use requires prior written authorization |
The MiniMax model catalog, API overview, and model release notes establish the status, context limit, model IDs, and release sequence above.
What MiniMax M2.7 is designed to do
MiniMax presents M2.7 as a model for complex agent harnesses and professional productivity tasks. Its official repository highlights software engineering work such as log analysis, bug investigation, refactoring, code-security review, machine-learning tasks, terminal operation, and multi-step tool use. It also describes Word, Excel, and presentation workflows that require editable deliverables and repeated revision.
The phrase “recursive self-improvement” in MiniMax’s release materials refers to how an internal version of M2.7 participated in parts of the model-development process. MiniMax says the system updated memory, built skills for reinforcement-learning experiments, analyzed failures, modified a programming scaffold, ran evaluations, and decided whether to keep or revert changes. It does not mean that every API request retrains the deployed model or allows it to alter its own production weights.
M2.7 status versus MiniMax M3 and legacy models
| Model family | Status verified July 17, 2026 | Combined context | Compatible content input |
|---|---|---|---|
| MiniMax-M3 | Frontier M-series model; released June 1, 2026 | 1,000,000 tokens | Text, images, video, and tool content; text output |
| MiniMax-M2.7 / M2.7-highspeed | Listed outside Legacy Models; released March 18, 2026 | 204,800 tokens | Text and tool-call content |
| M2.5, M2.1, and M2 variants | Listed under Legacy Models | 204,800 tokens in the API overview | Text and tool-call content |
For a new evaluation, test MiniMax M3 first when the workflow needs image or video input, a substantially larger context allowance, or the provider-designated frontier model. M2.7 remains relevant when a team has an established M2.7 integration, needs to reproduce an M2.7 evaluation, prefers its price profile, or plans a permitted self-hosted deployment using the public weights.
Do not infer that an older model is unavailable merely because M3 was released. Conversely, do not infer that inclusion in an API compatibility table makes M2.5, M2.1, or M2 equal in status to M2.7; the dedicated catalog labels those families as legacy.
Context window: what 204,800 tokens means
MiniMax documents a 204,800-token context window for both M2.7 variants. The provider explicitly defines this as the combined total of input and output tokens. It is not a promise that an application can always send 204,800 input tokens and still request a large completion.
Plan the token budget as:
system instructions + conversation history + tool definitions + tool results + user input + model reasoning and answer ≤ 204,800 tokens
Long-context capacity does not ensure that every detail receives equal attention. For repository analysis or document workflows, split the task, retrieve only relevant material, label sources, and verify that the answer cites the supplied evidence. Monitor actual token usage rather than estimating from character count alone.
MiniMax M2.7 API pricing
The official pay-as-you-go table showed the following standard-tier rates when verified on July 17, 2026. Prices are in U.S. dollars per 1 million tokens.
| Model | Input | Output | Prompt-cache read | Prompt-cache write |
|---|---|---|---|---|
MiniMax-M2.7 | $0.30 | $1.20 | $0.06 | $0.375 |
MiniMax-M2.7-highspeed | $0.60 | $2.40 | $0.06 | $0.375 |
MiniMax documents Priority admission at 1.5 times the standard price. Priority affects request admission and reliability; it does not change the model’s context limit. Check the official pay-as-you-go pricing page before budgeting, and use billed token usage rather than a word-count estimate. Our separate MiniMax pricing guide explains the difference between pay-as-you-go keys and Token Plan subscription keys.
M2.7 versus M2.7-highspeed
MiniMax describes the highspeed variant as having the same performance profile with faster inference. The API overview gives approximate output speeds of 60 tokens per second for M2.7 and 100 tokens per second for M2.7-highspeed. Highspeed doubles the listed input and output token prices, while the two cache rates in the pricing table are identical.
Choose between them using an application-level test. Measure time to first token, output speed, end-to-end tool-loop duration, failure rate, cache hit rate, and total cost with your prompts. A provider figure is not a service-level guarantee, and tool latency may dominate model-generation time.
How to call MiniMax M2.7 through the hosted API
MiniMax supports OpenAI-compatible and Anthropic-compatible interfaces. Its documentation labels the Anthropic SDK route as recommended, while the OpenAI-compatible route is useful for applications built around Chat Completions. The OpenAI-compatible base URL is https://api.minimax.io/v1.
Node.js example with the OpenAI-compatible endpoint
const response = await fetch("https://api.minimax.io/v1/chat/completions", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.MINIMAX_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "MiniMax-M2.7",
messages: [
{
role: "system",
content: "You are a careful software-engineering assistant."
},
{
role: "user",
content: "Review this function and propose a tested patch."
}
],
max_completion_tokens: 2048,
temperature: 1,
top_p: 0.9,
reasoning_split: true
})
});
if (!response.ok) {
throw new Error(`MiniMax API error ${response.status}: ${await response.text()}`);
}
const data = await response.json();
console.log(data.choices[0].message.content);
Store the key on the server, never in browser JavaScript or a public repository. Add timeouts, retries with backoff, rate-limit handling, request identifiers, spend limits, and output validation before production use. See the official OpenAI-compatible documentation and our independent MiniMax API guide.
Reasoning and multi-turn tool calls
MiniMax documents that thinking cannot be disabled for M2.x models. With OpenAI-compatible tool conversations, preserve the complete assistant response—including tool calls and reasoning content—when you append it to conversation history. With the Anthropic-compatible format, preserve every content block, including thinking, text, and tool-use blocks. Removing those blocks can break reasoning continuity.
M2.7 accepts text and tool-call content through the compatible APIs. It does not accept image or video content blocks there. Use M3 if the workflow needs supported image or video input.
Hosted API versus self-hosting M2.7
| Decision area | MiniMax-hosted API | Self-hosted weights |
|---|---|---|
| Operations | MiniMax operates model serving; you operate the application and integration | You operate or contract for inference, scaling, updates, monitoring, and incident response |
| Commercial terms | Open Platform terms and API billing apply | The weight license applies; commercial use requires prior written MiniMax authorization |
| Data path | Prompts and outputs pass through the MiniMax service under the applicable terms and privacy policy | Data handling depends on your runtime, host, logs, telemetry, storage, backups, and access controls |
| Capacity | No model-weight infrastructure to provision; account limits and provider availability apply | Large public weights require substantial compute, storage, and inference engineering |
| Updates | Provider manages serving changes | You control versions and must validate every change |
The official MiniMax M2.7 repository links the weights and deployment guides for SGLang, vLLM, Transformers, and ModelScope. Self-hosting is not equivalent to calling the hosted API: behavior can differ with quantization, inference engine, kernels, prompt template, sampling settings, tool implementation, and available hardware.
MiniMax M2.7 license explained
The public weights are not distributed under an unrestricted MIT license. MiniMax’s repository labels the file a Non-Commercial License. It permits specified non-commercial uses under MIT-style terms, including personal self-hosting and non-commercial research or education. The license says commercial use requires separate prior written authorization from MiniMax and requires a prominent “Built with MiniMax M2.7” notice.
The license also contains prohibited-use conditions and an “as is” warranty disclaimer. Read the official M2.7 license before downloading, modifying, fine-tuning, serving, or embedding the weights in a product.
Do not assume the weight license governs a paid request sent to MiniMax’s hosted API. Hosted service use is governed by the Open Platform agreement and applicable service rules. Likewise, paying API fees does not grant permission to commercially self-host the weights.
Officially reported M2.7 benchmarks
MiniMax reports the following results in its official repository. These are provider-reported figures, not tests conducted by MiniMax-AI.chat. Benchmark versions, agent scaffolds, tool access, compute budgets, prompts, pass counts, and scoring rules can materially affect results.
| Benchmark | MiniMax-reported M2.7 result | What it evaluates broadly |
|---|---|---|
| SWE-Pro | 56.22% | Professional software-engineering tasks |
| SWE Multilingual | 76.5 | Software tasks across languages |
| Multi SWE Bench | 52.7 | Repository-level software work |
| VIBE-Pro | 55.6% | Real-world coding-agent work |
| Terminal Bench 2 | 57.0% | Terminal and environment interaction |
| Toolathon | 46.3% | Tool-use tasks |
| MM Claw end-to-end | 62.7% | Complex skills and agent workflows |
| GDPval-AA | 1495 ELO | Professional-work deliverables |
Use these numbers to select tests, not to declare a universal winner. Run M2.7 and M3 on the same private task set with the same tools, time limits, retry policy, review rubric, and cost accounting. Measure task completion, correctness, regressions, human correction time, latency, and total spend.
Practical use cases
- Repository maintenance: bug investigation, refactoring plans, test generation, log correlation, and patch review.
- Tool-using agents: multi-step workflows that call search, terminal, database, or internal application tools.
- Document workflows: drafting and revising structured reports, spreadsheets, and presentation content with human approval.
- Long-document analysis: source-grounded synthesis within a 204,800-token combined budget.
- Model research: permitted non-commercial evaluation of agent scaffolds, self-hosting, sampling, and inference trade-offs.
Human review remains necessary for security patches, financial or legal material, production commands, data changes, and decisions that can affect people. Tool access should follow least privilege, with sandboxing, approval gates, audit logs, and reversible actions.
Evaluation checklist before deployment
- Confirm that the exact model ID is available to the intended account and region.
- Compare M2.7 with M3 on the same task set.
- Budget input, output, reasoning, tool results, and cache operations.
- Preserve complete assistant reasoning and tool-call messages in multi-turn history.
- Validate every tool argument before execution.
- Block secrets and sensitive data that the service contract does not permit.
- Measure task success and human correction effort rather than relying only on benchmark scores.
- For self-hosting, obtain commercial authorization when applicable and satisfy the attribution requirement.
- Pin model, runtime, prompt-template, and evaluation versions.
Frequently asked questions
Is MiniMax M2.7 a legacy model?
No, not in the official catalog verified on July 17, 2026. M2.7 and M2.7-highspeed appear outside the Legacy Models section. M2.5, M2.1, M2, and their listed highspeed variants appear under Legacy Models.
Is MiniMax M2.7 the frontier M-series model?
No. MiniMax released M3 on June 1, 2026 and labels it the frontier multimodal coding model. M2.7 remains a documented alternative.
Does M2.7 have a 204,800-token input window?
The 204,800 figure is the combined maximum for input and output, not an input-only allowance. Reserve enough capacity for reasoning and the answer.
Can M2.7 accept images or video?
Not through the documented OpenAI- and Anthropic-compatible content formats. MiniMax documents text and tool-call blocks for M2.x; supported image and video input belongs to M3.
Can M2.7 thinking be disabled?
No. MiniMax’s compatible API documentation says thinking cannot be disabled for M2.x models. A disabled setting may be accepted, but M2.7 continues to use thinking.
Is MiniMax M2.7 open source under MIT?
No. The weights are publicly available, but the repository uses a non-commercial license with MIT-style permissions for specified non-commercial uses. Commercial use requires prior written authorization and the required “Built with MiniMax M2.7” notice.
Should a new project use M2.7 or M3?
Evaluate M3 first for the provider-designated frontier model, multimodal input, or a 1-million-token combined context. Evaluate M2.7 when its cost, established integration, benchmark profile, or permitted self-hosting path fits the project. Use a controlled test rather than the model name alone.
Conclusion
MiniMax M2.7 remains a documented, non-legacy language-model option for text, tool use, coding agents, and professional workflows. Its documented strengths are balanced by clear boundaries: a 204,800-token combined context, no image or video input through the compatible interfaces, mandatory M2.x thinking, and a restrictive license for commercial self-hosting.
Present M2.7 as an available predecessor to M3, not as the frontier M-series model. Separate hosted API terms from the public-weight license, verify prices and account access before deployment, and compare both M2.7 variants with M3 using the same tasks and controls.
