MiniMax M2.7: Specs, API Pricing, License, and M3 Status

MiniMax M2.7 is a text-and-tool-use model released by MiniMax on March 18, 2026 for software engineering, agent workflows, and professional document tasks. It remains listed in MiniMax’s main language-model catalog, but MiniMax-M3 has been the frontier M-series model since June 1, 2026.

Verified model status

Last verified: July 17, 2026. MiniMax’s official catalog lists MiniMax-M3, MiniMax-M2.7, and MiniMax-M2.7-highspeed outside the Legacy Models section. It places MiniMax-M2.5, M2.5-highspeed, M2.1, M2.1-highspeed, and M2 under Legacy Models. M2.7 is therefore still a documented, non-legacy option, but it should not be presented as MiniMax’s frontier M-series model.

This is an independent guide and is not affiliated with or endorsed by MiniMax. Product status, account access, prices, limits, and license terms can change; verify them with the linked MiniMax sources before deployment.

MiniMax M2.7 at a glance

Official model IDMiniMax-M2.7
Faster hosted variantMiniMax-M2.7-highspeed
Release dateMarch 18, 2026
Catalog status verified July 17, 2026Listed outside the Legacy Models section; M3 is the frontier M-series model
Context limit204,800 tokens combined across input and output
Compatible input typesText and tool-call content; no image or video input for M2.x through the compatible interfaces
Documented hosted speedApproximately 60 tokens/second for M2.7 and 100 tokens/second for highspeed; these are documentation figures, not a latency guarantee
Hosted accessMiniMax Open Platform API; MiniMax also links M2.7 from its Agent and Token Plan materials
Local accessPublic weights from MiniMax’s official repository and Hugging Face organization
Weight licenseNon-commercial license with MIT-style permissions for permitted non-commercial uses; commercial use requires prior written authorization

The MiniMax model catalog, API overview, and model release notes establish the status, context limit, model IDs, and release sequence above.

What MiniMax M2.7 is designed to do

MiniMax presents M2.7 as a model for complex agent harnesses and professional productivity tasks. Its official repository highlights software engineering work such as log analysis, bug investigation, refactoring, code-security review, machine-learning tasks, terminal operation, and multi-step tool use. It also describes Word, Excel, and presentation workflows that require editable deliverables and repeated revision.

The phrase “recursive self-improvement” in MiniMax’s release materials refers to how an internal version of M2.7 participated in parts of the model-development process. MiniMax says the system updated memory, built skills for reinforcement-learning experiments, analyzed failures, modified a programming scaffold, ran evaluations, and decided whether to keep or revert changes. It does not mean that every API request retrains the deployed model or allows it to alter its own production weights.

M2.7 status versus MiniMax M3 and legacy models

Model familyStatus verified July 17, 2026Combined contextCompatible content input
MiniMax-M3Frontier M-series model; released June 1, 20261,000,000 tokensText, images, video, and tool content; text output
MiniMax-M2.7 / M2.7-highspeedListed outside Legacy Models; released March 18, 2026204,800 tokensText and tool-call content
M2.5, M2.1, and M2 variantsListed under Legacy Models204,800 tokens in the API overviewText and tool-call content

For a new evaluation, test MiniMax M3 first when the workflow needs image or video input, a substantially larger context allowance, or the provider-designated frontier model. M2.7 remains relevant when a team has an established M2.7 integration, needs to reproduce an M2.7 evaluation, prefers its price profile, or plans a permitted self-hosted deployment using the public weights.

Do not infer that an older model is unavailable merely because M3 was released. Conversely, do not infer that inclusion in an API compatibility table makes M2.5, M2.1, or M2 equal in status to M2.7; the dedicated catalog labels those families as legacy.

Context window: what 204,800 tokens means

MiniMax documents a 204,800-token context window for both M2.7 variants. The provider explicitly defines this as the combined total of input and output tokens. It is not a promise that an application can always send 204,800 input tokens and still request a large completion.

Plan the token budget as:

system instructions + conversation history + tool definitions + tool results + user input + model reasoning and answer ≤ 204,800 tokens

Long-context capacity does not ensure that every detail receives equal attention. For repository analysis or document workflows, split the task, retrieve only relevant material, label sources, and verify that the answer cites the supplied evidence. Monitor actual token usage rather than estimating from character count alone.

MiniMax M2.7 API pricing

The official pay-as-you-go table showed the following standard-tier rates when verified on July 17, 2026. Prices are in U.S. dollars per 1 million tokens.

ModelInputOutputPrompt-cache readPrompt-cache write
MiniMax-M2.7$0.30$1.20$0.06$0.375
MiniMax-M2.7-highspeed$0.60$2.40$0.06$0.375

MiniMax documents Priority admission at 1.5 times the standard price. Priority affects request admission and reliability; it does not change the model’s context limit. Check the official pay-as-you-go pricing page before budgeting, and use billed token usage rather than a word-count estimate. Our separate MiniMax pricing guide explains the difference between pay-as-you-go keys and Token Plan subscription keys.

M2.7 versus M2.7-highspeed

MiniMax describes the highspeed variant as having the same performance profile with faster inference. The API overview gives approximate output speeds of 60 tokens per second for M2.7 and 100 tokens per second for M2.7-highspeed. Highspeed doubles the listed input and output token prices, while the two cache rates in the pricing table are identical.

Choose between them using an application-level test. Measure time to first token, output speed, end-to-end tool-loop duration, failure rate, cache hit rate, and total cost with your prompts. A provider figure is not a service-level guarantee, and tool latency may dominate model-generation time.

How to call MiniMax M2.7 through the hosted API

MiniMax supports OpenAI-compatible and Anthropic-compatible interfaces. Its documentation labels the Anthropic SDK route as recommended, while the OpenAI-compatible route is useful for applications built around Chat Completions. The OpenAI-compatible base URL is https://api.minimax.io/v1.

Node.js example with the OpenAI-compatible endpoint

const response = await fetch("https://api.minimax.io/v1/chat/completions", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.MINIMAX_API_KEY}`,
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    model: "MiniMax-M2.7",
    messages: [
      {
        role: "system",
        content: "You are a careful software-engineering assistant."
      },
      {
        role: "user",
        content: "Review this function and propose a tested patch."
      }
    ],
    max_completion_tokens: 2048,
    temperature: 1,
    top_p: 0.9,
    reasoning_split: true
  })
});

if (!response.ok) {
  throw new Error(`MiniMax API error ${response.status}: ${await response.text()}`);
}

const data = await response.json();
console.log(data.choices[0].message.content);

Store the key on the server, never in browser JavaScript or a public repository. Add timeouts, retries with backoff, rate-limit handling, request identifiers, spend limits, and output validation before production use. See the official OpenAI-compatible documentation and our independent MiniMax API guide.

Reasoning and multi-turn tool calls

MiniMax documents that thinking cannot be disabled for M2.x models. With OpenAI-compatible tool conversations, preserve the complete assistant response—including tool calls and reasoning content—when you append it to conversation history. With the Anthropic-compatible format, preserve every content block, including thinking, text, and tool-use blocks. Removing those blocks can break reasoning continuity.

M2.7 accepts text and tool-call content through the compatible APIs. It does not accept image or video content blocks there. Use M3 if the workflow needs supported image or video input.

Hosted API versus self-hosting M2.7

Decision areaMiniMax-hosted APISelf-hosted weights
OperationsMiniMax operates model serving; you operate the application and integrationYou operate or contract for inference, scaling, updates, monitoring, and incident response
Commercial termsOpen Platform terms and API billing applyThe weight license applies; commercial use requires prior written MiniMax authorization
Data pathPrompts and outputs pass through the MiniMax service under the applicable terms and privacy policyData handling depends on your runtime, host, logs, telemetry, storage, backups, and access controls
CapacityNo model-weight infrastructure to provision; account limits and provider availability applyLarge public weights require substantial compute, storage, and inference engineering
UpdatesProvider manages serving changesYou control versions and must validate every change

The official MiniMax M2.7 repository links the weights and deployment guides for SGLang, vLLM, Transformers, and ModelScope. Self-hosting is not equivalent to calling the hosted API: behavior can differ with quantization, inference engine, kernels, prompt template, sampling settings, tool implementation, and available hardware.

MiniMax M2.7 license explained

The public weights are not distributed under an unrestricted MIT license. MiniMax’s repository labels the file a Non-Commercial License. It permits specified non-commercial uses under MIT-style terms, including personal self-hosting and non-commercial research or education. The license says commercial use requires separate prior written authorization from MiniMax and requires a prominent “Built with MiniMax M2.7” notice.

The license also contains prohibited-use conditions and an “as is” warranty disclaimer. Read the official M2.7 license before downloading, modifying, fine-tuning, serving, or embedding the weights in a product.

Do not assume the weight license governs a paid request sent to MiniMax’s hosted API. Hosted service use is governed by the Open Platform agreement and applicable service rules. Likewise, paying API fees does not grant permission to commercially self-host the weights.

Officially reported M2.7 benchmarks

MiniMax reports the following results in its official repository. These are provider-reported figures, not tests conducted by MiniMax-AI.chat. Benchmark versions, agent scaffolds, tool access, compute budgets, prompts, pass counts, and scoring rules can materially affect results.

BenchmarkMiniMax-reported M2.7 resultWhat it evaluates broadly
SWE-Pro56.22%Professional software-engineering tasks
SWE Multilingual76.5Software tasks across languages
Multi SWE Bench52.7Repository-level software work
VIBE-Pro55.6%Real-world coding-agent work
Terminal Bench 257.0%Terminal and environment interaction
Toolathon46.3%Tool-use tasks
MM Claw end-to-end62.7%Complex skills and agent workflows
GDPval-AA1495 ELOProfessional-work deliverables

Use these numbers to select tests, not to declare a universal winner. Run M2.7 and M3 on the same private task set with the same tools, time limits, retry policy, review rubric, and cost accounting. Measure task completion, correctness, regressions, human correction time, latency, and total spend.

Practical use cases

  • Repository maintenance: bug investigation, refactoring plans, test generation, log correlation, and patch review.
  • Tool-using agents: multi-step workflows that call search, terminal, database, or internal application tools.
  • Document workflows: drafting and revising structured reports, spreadsheets, and presentation content with human approval.
  • Long-document analysis: source-grounded synthesis within a 204,800-token combined budget.
  • Model research: permitted non-commercial evaluation of agent scaffolds, self-hosting, sampling, and inference trade-offs.

Human review remains necessary for security patches, financial or legal material, production commands, data changes, and decisions that can affect people. Tool access should follow least privilege, with sandboxing, approval gates, audit logs, and reversible actions.

Evaluation checklist before deployment

  • Confirm that the exact model ID is available to the intended account and region.
  • Compare M2.7 with M3 on the same task set.
  • Budget input, output, reasoning, tool results, and cache operations.
  • Preserve complete assistant reasoning and tool-call messages in multi-turn history.
  • Validate every tool argument before execution.
  • Block secrets and sensitive data that the service contract does not permit.
  • Measure task success and human correction effort rather than relying only on benchmark scores.
  • For self-hosting, obtain commercial authorization when applicable and satisfy the attribution requirement.
  • Pin model, runtime, prompt-template, and evaluation versions.

Frequently asked questions

Is MiniMax M2.7 a legacy model?

No, not in the official catalog verified on July 17, 2026. M2.7 and M2.7-highspeed appear outside the Legacy Models section. M2.5, M2.1, M2, and their listed highspeed variants appear under Legacy Models.

Is MiniMax M2.7 the frontier M-series model?

No. MiniMax released M3 on June 1, 2026 and labels it the frontier multimodal coding model. M2.7 remains a documented alternative.

Does M2.7 have a 204,800-token input window?

The 204,800 figure is the combined maximum for input and output, not an input-only allowance. Reserve enough capacity for reasoning and the answer.

Can M2.7 accept images or video?

Not through the documented OpenAI- and Anthropic-compatible content formats. MiniMax documents text and tool-call blocks for M2.x; supported image and video input belongs to M3.

Can M2.7 thinking be disabled?

No. MiniMax’s compatible API documentation says thinking cannot be disabled for M2.x models. A disabled setting may be accepted, but M2.7 continues to use thinking.

Is MiniMax M2.7 open source under MIT?

No. The weights are publicly available, but the repository uses a non-commercial license with MIT-style permissions for specified non-commercial uses. Commercial use requires prior written authorization and the required “Built with MiniMax M2.7” notice.

Should a new project use M2.7 or M3?

Evaluate M3 first for the provider-designated frontier model, multimodal input, or a 1-million-token combined context. Evaluate M2.7 when its cost, established integration, benchmark profile, or permitted self-hosting path fits the project. Use a controlled test rather than the model name alone.

Conclusion

MiniMax M2.7 remains a documented, non-legacy language-model option for text, tool use, coding agents, and professional workflows. Its documented strengths are balanced by clear boundaries: a 204,800-token combined context, no image or video input through the compatible interfaces, mandatory M2.x thinking, and a restrictive license for commercial self-hosting.

Present M2.7 as an available predecessor to M3, not as the frontier M-series model. Separate hosted API terms from the public-weight license, verify prices and account access before deployment, and compare both M2.7 variants with M3 using the same tasks and controls.