MiniMax-Text-01 is a long-context language model from the MiniMax-01 series. It has 456B total parameters, activates 45.9B parameters per token, uses a hybrid attention design with Mixture-of-Experts, and is described as supporting up to a 4M-token inference context in the MiniMax-01 release materials.
Last updated: June 2026. MiniMax-Text-01 remains important as an open-weight MiniMax-01 long-context model, but MiniMax’s current API documentation now lists MiniMax-M3 as the latest M-series model and focuses the main supported-model tables on M3 and M2-series routes. Verify live first-party API availability before building production workflows around MiniMax-Text-01.
What Is MiniMax-Text-01?
MiniMax-Text-01 is the text-focused language model in the MiniMax-01 series. It is designed for text generation, reasoning, long-document processing, tool use, and developer workflows that need unusually long context handling.
The MiniMax-01 series includes two named models: MiniMax-Text-01 for language tasks and MiniMax-VL-01 for vision-language tasks. The arXiv technical report describes MiniMax-01 as a series that includes both MiniMax-Text-01 and MiniMax-VL-01, with the text model serving as the foundation for the vision-language variant.
For searchers comparing long-context LLMs, the most important point is that MiniMax-Text-01 is not just a model name. It refers to a specific architecture: a large Mixture-of-Experts language model that combines Lightning Attention, Softmax Attention, and MoE routing to scale context length while controlling per-token compute.
MiniMax-Text-01 Specifications
| Specification | Detail |
|---|---|
| Model family | MiniMax-01 |
| Model type | Long-context language model |
| Developer | MiniMax AI |
| Release period | January 2025 |
| Total parameters | 456B |
| Activated parameters per token | 45.9B |
| Architecture | Lightning Attention + Softmax Attention + Mixture-of-Experts |
| Layers | 80 |
| Experts | 32 |
| Routing | Top-2 MoE routing |
| Hidden size | 6144 |
| Vocabulary size | 200,064 |
| Training context length | 1M tokens |
| Inference context length | Up to 4M tokens as reported by the model card / technical report; this is a release-material inference claim and should not be assumed for every hosted API or third-party provider route. |
| Current MiniMax API docs status | MiniMax’s current Model Invocation docs list MiniMax-M3 as the latest M-series model and focus the main supported-model table on M3 and M2-series routes. Verify live API availability before building first-party production workflows around MiniMax-Text-01. |
| Function calling | Supported |
| Deployment ecosystem | Hugging Face / Transformers quickstart and MiniMax’s own vLLM deployment guide exist; verify current serving-framework and provider support before production use. |
According to the MiniMax model card and GitHub repository, MiniMax-Text-01 has 456B total parameters, 45.9B activated parameters per token, 80 layers, 32 experts, Top-2 routing, a 6144 hidden size, and a 200,064-token vocabulary. The same materials describe a 1M-token training context and up to 4M tokens during inference.
How MiniMax-Text-01 Fits Into the MiniMax-01 Series
The naming can be confusing because MiniMax-01 may refer to the broader model series, while MiniMax-Text-01 refers to the text model inside that series.
MiniMax-Text-01 is the language model. It is built for text generation, reasoning, long-context processing, code-related tasks, structured outputs, and tool-use workflows.
MiniMax-VL-01 is the vision-language model. The GitHub repository describes it as being built on MiniMax-Text-01 with additional visual components, including a 303M-parameter Vision Transformer, an MLP projector, and a dynamic resolution mechanism for image inputs.
MiniMax-01 is the family or release name. In practical discussions, people may use “MiniMax-01” loosely, but a technical article should separate the series name from the specific text and multimodal models.
Architecture: Lightning Attention, Softmax Attention, and MoE
MiniMax-Text-01 uses a hybrid architecture that combines Lightning Attention, Softmax Attention, and Mixture-of-Experts. This design is central to why the model is associated with long-context use cases.
Lightning Attention is used to improve long-sequence efficiency. Traditional attention becomes expensive as sequence length grows, so long-context models often need architectural changes to process very large prompts without prohibitive cost.
Softmax Attention remains part of the architecture because full attention is still useful for preserving strong modeling quality. The MiniMax model card describes the hybrid attention layout as placing one Softmax Attention layer after every seven Lightning Attention layers.
Mixture-of-Experts allows the model to have a very large total parameter count while activating only a subset of parameters for each token. In MiniMax-Text-01, the reported total size is 456B parameters, while 45.9B parameters are activated per token. This helps explain how the model can combine scale with more controlled inference compute than a dense model of the same total size.
The model card also reports 32 experts and a Top-2 routing strategy. In simple terms, Top-2 routing means the model selects two expert pathways for a token rather than sending every token through every expert.
Why the 4M-Token Context Window Matters
The context window is one of the main reasons MiniMax-Text-01 attracts attention. The MiniMax-01 technical report states that the context window can reach 1M tokens during training and extrapolate to 4M tokens during inference.
That distinction is important. A 1M-token training context means the model was trained to handle very long sequences. A reported 4M-token inference context means the model can be used for even larger input windows under the evaluation and serving assumptions described in the release materials.
However, developers should not assume every hosted API or third-party provider exposes the full 4M-token limit. Provider-specific limits may vary, and production constraints such as cost, latency, rate limits, memory, and safety filters can affect how much context is practical in a real application.
A 4M-token context can matter for tasks such as:
- Long-document analysis: processing books, reports, transcripts, manuals, or policy documents.
- Large codebases: reading many files together instead of analyzing isolated snippets.
- Legal or financial document collections: comparing contracts, filings, disclosures, or evidence bundles.
- Research papers: synthesizing many papers, appendices, tables, and methodology notes.
- Multi-agent memory: giving agent systems a longer working history across tasks.
- Retrieval-augmented generation: combining retrieved passages with broader project context.
Long context is not a replacement for good retrieval design, but it can reduce the amount of aggressive truncation required when an application needs to preserve a large amount of source material.
Benchmarks and Reported Performance
The MiniMax model card reports benchmark scores across academic, coding, instruction-following, and long-context evaluations. These are reported benchmark results from the MiniMax model card / technical report and should be interpreted within the evaluation setup used by the authors.
| Benchmark | Reported MiniMax-Text-01 Score |
|---|---|
| MMLU | 88.5 |
| MMLU-Pro | 75.7 |
| IFEval average | 89.1 |
| Arena-Hard | 89.1 |
| HumanEval | 86.9 |
| LongBench v2 with CoT overall | 56.5 |
| RULER at 1M | 0.910 |
The model card reports MMLU at 88.5, MMLU-Pro at 75.7, IFEval average at 89.1, Arena-Hard at 89.1, and HumanEval at 86.9. For long-context evaluation, it reports a RULER score of 0.910 at 1M tokens and a LongBench v2 with chain-of-thought overall score of 56.5.
These numbers are useful for comparison, but they should not be treated as a guarantee of application performance. Benchmarks depend on prompt format, evaluation settings, decoding parameters, data contamination controls, and task fit. A model that performs strongly on a benchmark may still need careful testing in a domain-specific workflow.
MiniMax-Text-01 API and Developer Use
Developers may encounter MiniMax-Text-01 through self-hosting, Hugging Face / Transformers, MiniMax launch-time API references, or third-party provider routes. The model card includes a Transformers quickstart, MiniMax’s own vLLM deployment guide, and a function-calling section. However, MiniMax’s current Model Invocation docs list MiniMax-M3 and M2-series models in the main supported-model table, so developers should verify live first-party API availability before building around MiniMax-Text-01.
Function calling is one of the model’s developer-oriented capabilities. The model card says MiniMax-Text-01 can identify when external functions need to be called and output structured JSON-format parameters. It also notes support for complex parameter types, including nested objects and arrays.
Below is a generic Python-style example showing how a developer might structure an API call. This is illustrative only; endpoint names, authentication, request fields, pricing, and context limits depend on the provider.
import os
import requests
API_KEY = os.getenv("MINIMAX_API_KEY")
API_URL = "https://api.example-provider.com/v1/chat/completions"
payload = {
"model": "MiniMax-Text-01",
"messages": [
{
"role": "system",
"content": "You are a technical assistant that summarizes long documents."
},
{
"role": "user",
"content": "Summarize the attached research notes and identify open questions."
}
],
"temperature": 0.3,
"tools": [
{
"type": "function",
"function": {
"name": "save_summary",
"description": "Save a structured summary for later review.",
"parameters": {
"type": "object",
"properties": {
"title": {"type": "string"},
"key_points": {
"type": "array",
"items": {"type": "string"}
},
"open_questions": {
"type": "array",
"items": {"type": "string"}
}
},
"required": ["title", "key_points"]
}
}
}
]
}
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json"
}
response = requests.post(API_URL, headers=headers, json=payload, timeout=120)
response.raise_for_status()
print(response.json())
For production use, teams should verify the exact model ID, provider availability, context limit, rate policy, supported tool format, pricing, and data-handling terms before building around a hosted endpoint. Do not assume that the launch-time MiniMax-Text-01 API route is still available in the same way in the current MiniMax API docs.
Best Use Cases for MiniMax-Text-01
MiniMax-Text-01 is most relevant when a workflow benefits from both large model capacity and very long context handling.
- Long-document summarization: Summarize lengthy reports, books, transcripts, manuals, or policy archives while preserving more surrounding context.
- Document question answering: Ask detailed questions across large document sets, especially when answers depend on references spread across many pages.
- Agent memory and multi-step workflows: Give AI agents a longer memory trace across planning, execution, tool calls, and user feedback.
- Codebase analysis: Analyze multiple files, architecture notes, logs, and documentation together instead of relying only on small snippets.
- Research synthesis: Compare papers, technical reports, benchmark sections, appendices, and experimental notes in a unified prompt.
- Structured tool-use workflows: Use function calling to produce structured parameters for external systems, APIs, databases, or automation tools.
- Large-scale RAG pipelines: Combine retrieval with a larger working context so the model can reason over more retrieved evidence at once.
The best results usually come from pairing long context with strong prompt design, source ranking, evaluation datasets, and domain-specific validation.
Limitations and Practical Considerations
MiniMax-Text-01 is a very large model. Even though MoE routing reduces the number of active parameters per token, infrastructure requirements can still be significant. Teams considering self-hosting should account for memory, parallelism, quantization, serving throughput, and operational cost.
The full reported 4M-token context may not be exposed through every provider. Some providers may set lower input limits, lower output limits, different pricing tiers, or different batching rules. For production planning, provider-specific limits may vary.
Long context also does not remove the need for retrieval, chunking, ranking, or evaluation. A larger prompt can carry more material, but applications still need to decide what information is relevant, how to order it, and how to test whether the model used it correctly.
Benchmark scores should be treated as signals, not universal proof. Real-world performance depends on task type, prompt structure, output requirements, domain vocabulary, latency constraints, and failure tolerance.
Licensing should be reviewed carefully before commercial redistribution. The GitHub repository distinguishes repository licensing from model-material licensing, and the model license includes a grant of rights plus redistribution conditions. Avoid describing the model materials simply as “MIT” without checking the relevant license terms.
MiniMax-Text-01 vs Similar Long-Context Models
MiniMax-Text-01 is often discussed alongside models such as GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Llama 3.1 405B because the MiniMax-01 model card used those systems as launch-time comparison points across academic and long-context benchmarks. For current MiniMax-family model selection, also compare MiniMax-Text-01 with newer pages such as MiniMax M1, MiniMax M2.7, and MiniMax M3, because current API availability and model positioning have changed since the January 2025 release.
| Model | Comparison framing |
|---|---|
| MiniMax-Text-01 | Strong release-material positioning on long-context tasks, with 456B total parameters and 45.9B active parameters per token. |
| GPT-4o | Included in the MiniMax model card’s benchmark comparisons across general, coding, and long-context tasks. |
| Claude 3.5 Sonnet | Included as a high-performing comparison model in the MiniMax-01 materials. |
| Gemini 1.5 Pro | Particularly relevant in long-context comparisons because it appears in RULER and other long-context tables. |
| Llama 3.1 405B | Useful as an open-weight scale comparison in benchmark tables, though architecture and serving assumptions differ. |
Reported comparisons in the MiniMax-01 materials positioned MiniMax-Text-01 strongly on long-context tasks, while other models may perform better on specific reasoning, coding, latency, tool-use, or ecosystem criteria. A fair comparison should use the same prompts, data, evaluation criteria, and serving constraints.
Who Should Consider MiniMax-Text-01?
MiniMax-Text-01 is worth considering for teams that need to experiment with long-context language modeling at serious scale.
It is especially relevant for AI application developers building document-heavy products, RAG system builders testing larger context windows, researchers studying long-context architectures, and teams working with agent memory or multi-step workflows.
It may also interest technical founders evaluating whether long context can simplify parts of their product architecture. For example, a legal AI product, research assistant, code review system, or enterprise knowledge assistant may benefit from a model that can preserve more source material in a single interaction.
The model is less suitable for teams that need only short chat responses, lightweight inference, minimal infrastructure, or very low-latency responses on small prompts. In those cases, smaller models or hosted APIs with simpler deployment paths may be more practical.
Conclusion
MiniMax-Text-01 is important because it combines long-context design, hybrid attention, Mixture-of-Experts scale, and developer-oriented capabilities such as function calling. Its reported 456B total parameters, 45.9B active parameters per token, 1M-token training context, and up to 4M-token inference context make it a notable model for long-document workflows and AI-agent research.
For developers and researchers, the best way to evaluate MiniMax-Text-01 is to test it against the exact workload that matters: long documents, large codebases, structured tool calls, retrieval pipelines, or multi-step reasoning workflows. Benchmarks are useful starting points, but production value depends on task fit, provider constraints, cost, latency, and careful evaluation. For current MiniMax API production choices, compare against newer MiniMax M-series routes such as MiniMax M3, and for reasoning-focused open-weight workflows compare against MiniMax M1.
FAQ
What is MiniMax-Text-01?
MiniMax-Text-01 is a long-context language model in the MiniMax-01 series. It is designed for text generation, reasoning, long-document processing, and developer workflows that benefit from very large context windows.
Is MiniMax-Text-01 the same as MiniMax-01?
No. MiniMax-Text-01 is the text-focused model. MiniMax-01 is the broader series name that includes MiniMax-Text-01 and MiniMax-VL-01.
How many parameters does MiniMax-Text-01 have?
The MiniMax model card reports 456B total parameters, with 45.9B parameters activated per token.
What is the context window of MiniMax-Text-01?
The MiniMax-01 technical report and model card describe a 1M-token training context and up to 4M tokens during inference. Provider-specific limits may vary.
Does MiniMax-Text-01 support function calling?
Yes. The model card states that MiniMax-Text-01 supports function calling and can output structured JSON-format parameters for external function calls.
Can developers use MiniMax-Text-01 through an API?
The model card points to MiniMax launch-time online API access and includes local/deployment references such as Transformers and vLLM. However, MiniMax’s current Model Invocation docs focus on MiniMax-M3 and M2-series routes, so treat MiniMax-Text-01 API access as historical or provider-specific unless your live provider documentation lists it. Exact access, limits, and pricing depend on the provider.
Is MiniMax-Text-01 open-source?
MiniMax described the MiniMax-01 series as publicly released/open-sourced, but the safest wording is that MiniMax-Text-01 has public/open-weight model materials under the MiniMax Model License. The model license grants broad rights to use, reproduce, distribute, modify, and create derivative works, but it also includes redistribution, attribution, legal-compliance, and prohibited-use conditions. Teams should review the exact model license before redistribution, commercial use, derivative services, or fine-tuning.
What is MiniMax-Text-01 best used for?
MiniMax-Text-01 is best suited for long-document summarization, document question answering, codebase analysis, research synthesis, agent memory, structured tool use, and large-scale retrieval-augmented generation.
How is MiniMax-Text-01 different from MiniMax-VL-01?
MiniMax-Text-01 is the language model for text tasks. MiniMax-VL-01 is the vision-language model built with additional visual components on top of MiniMax-Text-01.
Does a 4M-token context window mean better answers automatically?
No. A larger context window can let the model process more material, but answer quality still depends on prompt design, retrieval quality, document ordering, task complexity, and evaluation.
