Last verified: August 13, 2026. Music API availability rechecked: August 20, 2026.
The MiniMax API gives developers server-side access to MiniMax language, speech, video, image, music, voice, and file services. For a new text integration, use MiniMax-M3 with either the OpenAI-compatible base URL https://api.minimax.io/v1 or the Anthropic-compatible base URL https://api.minimax.io/anthropic.
Independent guide: MiniMax-AI.chat is not the official MiniMax website, does not issue API keys, and cannot access your account. This page is a dated, docs-verified reference. Confirm the schema for the exact endpoint you use before production deployment.
MiniMax API Quick Reference
| What you need | Current value |
|---|---|
| International OpenAI-compatible base URL | https://api.minimax.io/v1 |
| International Anthropic-compatible base URL | https://api.minimax.io/anthropic |
| OpenAI Chat endpoint | POST /v1/chat/completions |
| Responses endpoint | POST /v1/responses |
| Anthropic Messages endpoint | POST /anthropic/v1/messages |
| Authentication | Authorization: Bearer YOUR_API_KEY |
| Current general-purpose model | MiniMax-M3 |
| China-region OpenAI / Anthropic bases | https://api.minimaxi.com/v1 / https://api.minimaxi.com/anthropic |
The base URL belongs in your SDK client configuration. Do not append /v1 twice. With the Anthropic SDK, set the base URL to https://api.minimax.io/anthropic; the SDK adds the Messages route.
What This MiniMax API Guide Covers
- How to get and store a MiniMax API key
- Working cURL, Python, Node.js, and Anthropic examples
- OpenAI Chat vs Responses vs Anthropic Messages
- Current model IDs and context limits
- Parameters, streaming, tools, and multimodal input
- Text, speech, video, image, music, voice, and file endpoints
- Pricing and a worked cost example
- Rate limits and common API errors
- Security and production checklist
- Verification scope and official sources
- MiniMax API FAQ
How to Get a MiniMax API Key
Create the credential in MiniMax’s official Open Platform after registration. The platform separates a pay-as-you-go API Key from the Subscription Key used by eligible Token Plan quota and Credits. Use the key that matches the billing resource you intend to spend.
- Open the official prerequisites guide and sign in to the correct regional platform.
- Create a pay-as-you-go API Key or the appropriate Subscription Key.
- Copy it once into a secret manager or a local environment variable.
- Never place the key in browser-side JavaScript, a mobile bundle, a public repository, a screenshot, or a support post.
- Use separate development and production credentials so usage and incident response remain attributable.
| Credential | Billing source | Best fit |
|---|---|---|
| Pay-as-you-go API Key | Open Platform wallet for the account or team | Metered development and production requests |
| Subscription Key | Eligible Token Plan quota, then eligible Credits | Token Plan seats and prepaid Credits |
See the focused guides to MiniMax API keys and base URLs and API Key vs Subscription Key for setup and billing details.
Store the key safely
# macOS or Linux
export MINIMAX_API_KEY="replace-with-your-server-side-key"
# PowerShell
$env:MINIMAX_API_KEY="replace-with-your-server-side-key"
Do not commit a real .env file. Add it to .gitignore and use your hosting provider’s secret-management feature in production.
Make Your First MiniMax API Call in Five Minutes
The examples below use the international OpenAI-compatible route and MiniMax-M3. They intentionally request a short answer. Replace the prompt, output limit, and error handling for your application.
cURL: send a real Chat Completions request
curl https://api.minimax.io/v1/chat/completions \
-H "Authorization: Bearer $MINIMAX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "MiniMax-M3",
"messages": [
{
"role": "user",
"content": "Reply with one short sentence confirming the API is connected."
}
],
"max_completion_tokens": 128,
"thinking": {
"type": "disabled"
}
}'
A successful response normally includes an assistant message plus token-usage data. Validate the HTTP status, the expected response object, and any MiniMax base_resp status fields returned by the specific endpoint. HTTP 200 alone should not be your only application-level success check.
Python with the OpenAI SDK
Install the client with pip install openai, then run:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.minimax.io/v1",
api_key=os.environ["MINIMAX_API_KEY"],
)
response = client.chat.completions.create(
model="MiniMax-M3",
messages=[
{
"role": "user",
"content": "Reply with one short sentence confirming the API is connected.",
}
],
max_completion_tokens=128,
extra_body={"thinking": {"type": "disabled"}},
)
print(response.choices[0].message.content)
print(response.usage)
Use max_completion_tokens for OpenAI-compatible Chat Completions. The older max_tokens parameter is deprecated on that interface.
Node.js with the OpenAI SDK
Install the client with npm install openai. Keep this code on your server, not in frontend JavaScript:
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.minimax.io/v1",
apiKey: process.env.MINIMAX_API_KEY,
});
const response = await client.chat.completions.create({
model: "MiniMax-M3",
messages: [
{
role: "user",
content: "Reply with one short sentence confirming the API is connected.",
},
],
max_completion_tokens: 128,
});
console.log(response.choices[0].message.content);
console.log(response.usage);
Python with the Anthropic SDK
MiniMax’s current quick-start documentation recommends the Anthropic SDK route. Install it with pip install anthropic:
import os
import anthropic
client = anthropic.Anthropic(
base_url="https://api.minimax.io/anthropic",
api_key=os.environ["MINIMAX_API_KEY"],
)
message = client.messages.create(
model="MiniMax-M3",
max_tokens=128,
messages=[
{
"role": "user",
"content": "Reply with one short sentence confirming the API is connected.",
}
],
)
for block in message.content:
if block.type == "text":
print(block.text)
Do not mechanically rename this parameter: max_tokens remains correct for Anthropic Messages. The three current output-limit names are max_completion_tokens for OpenAI Chat, max_output_tokens for Responses, and max_tokens for Anthropic Messages.
Which MiniMax API Interface Should You Use?
| Interface | Route | Output-limit field | Choose it when |
|---|---|---|---|
| OpenAI Chat Completions | POST /v1/chat/completions | max_completion_tokens | You already use the OpenAI SDK or Chat Completions message format. |
| OpenAI Responses | POST /v1/responses | max_output_tokens | Your application is built around the Responses object and its event model. |
| Anthropic Messages | POST /anthropic/v1/messages | max_tokens | You use the Anthropic SDK, Messages format, or supported thinking and tool workflows. |
Compatibility does not mean every upstream parameter or response field behaves identically. Validate the fields your application uses against the exact MiniMax operation reference. Start with our OpenAI-compatible API guide, Anthropic-compatible API guide, or Responses API guide.
Legacy migration note:
POST /v1/text/chatcompletion_v2is still documented but marked Deprecated. Do not use it as the starting point for a new integration; migrate maintained text applications to Chat Completions, Responses, or Anthropic Messages.
Current MiniMax API Models
The official model catalog listed the following language models when this page was verified. Context figures combine input and output; they do not guarantee that every client, plan, or request will expose the maximum.
| Model ID | Published context | Inputs | Status and use |
|---|---|---|---|
MiniMax-M3 | 1,000,000 tokens | Text, images, video, tools | Current general-purpose model for coding, agents, tools, and long-context work. |
MiniMax-M2.7 | 204,800 tokens | Text and tools | Active text model for engineering and professional workflows. |
MiniMax-M2.7-highspeed | 204,800 tokens | Text and tools | Speed-focused hosted M2.7 option. |
MiniMax-M2.5 variants | 204,800 tokens | Text and tools | Legacy; preserve only where an existing integration requires them. |
MiniMax-M2.1 variants and MiniMax-M2 | 204,800 tokens | Text and tools | Legacy. |
For MiniMax-M3, MiniMax documents a recommended output limit of 131,072 tokens and a hard maximum of 524,288. For other listed language models, the documented recommendation is 65,536 with a 204,800 maximum. Set a smaller explicit limit unless a tested workload needs more, and leave enough room for the input inside the total context budget.
See the detailed MiniMax M3 guide and MiniMax M2.7 guide. Hosted API availability and downloadable model-weight licensing are separate questions.
Parameters, Streaming, Tools, and Multimodal Input
- Streaming:
stream: trueis documented for Chat Completions, Responses, and Messages. Your client must process the event format for the selected interface. - Tools: use the modern
toolsfield. The olderfunction_callfield is not supported.tool_choicesupportsautoandnone. - One output per request: the Chat Completions
nfield supports1. - Tool history: in a multi-turn tool workflow, retain and resend the complete assistant response, including tool calls and supported thinking fields.
- Multimodal:
MiniMax-M3accepts text, image, and video input on documented routes. Audio input is not supported by Chat Completions. M2.x models are text-and-tools models. - Thinking: thinking behavior and defaults differ by interface. Configure it deliberately and preserve the returned state when the workflow requires it.
For implementation details, use the dedicated guides to MiniMax function calling, prompt caching, M3 multimodal input, and Anthropic Server Tools web search. Server Tools is currently a beta Anthropic Messages feature; the documented web_search tool costs $0.01 per request.
MiniMax API Endpoint Map
All routes below use https://api.minimax.io as the international host. A chat endpoint does not generate every media type: text, speech, video, image, music, voice, and managed files use separate schemas.
Language endpoints
| Capability | Method and route |
|---|---|
| List OpenAI-compatible models | GET /v1/models |
| Chat Completions | POST /v1/chat/completions |
| Responses | POST /v1/responses |
| Count Responses input tokens | POST /v1/responses/input_tokens |
| Anthropic Messages | POST /anthropic/v1/messages |
| Count Anthropic input tokens | POST /anthropic/v1/messages/count_tokens |
Media, voice, and file endpoints
| Capability | Method and route | Workflow |
|---|---|---|
| Synchronous text to speech | POST /v1/t2a_v2 | Return speech over HTTP; a WebSocket interface is also documented. |
| Asynchronous long speech | POST /v1/t2a_async_v2 | Create a task, then poll /v1/query/t2a_async_query_v2. |
Video Generation V2 (MiniMax-H3) | POST /v2/video_generation | Create a task, then query GET /v2/query/video_generation/{task_id}; Pay as You Go is required. |
| Legacy Hailuo video generation | POST /v1/video_generation | Create a task, then poll GET /v1/query/video_generation?task_id=.... |
| Image generation | POST /v1/image_generation | Text-to-image or supported subject-reference workflow. |
| Music generation | POST /v1/music_generation | Song, instrumental, or supported cover workflow. |
| Lyrics generation | POST /v1/lyrics_generation | Create, edit, or continue lyrics. |
| Voice cloning | POST /v1/voice_clone | Create a voice from eligible uploaded audio. |
| Voice design | POST /v1/voice_design | Design a voice from a description. |
| Upload a managed file | POST /v1/files/upload | Upload with the purpose required by the destination endpoint. |
| Retrieve file metadata | GET /v1/files/retrieve?file_id=... | Read metadata or a generated-file URL. |
| Retrieve file content | GET /v1/files/retrieve_content?file_id=... | Download managed content where supported. |
The model catalog verified on August 13, 2026 lists MiniMax-H3 as the current video model; it uses the V2 API. Hailuo 2.3, Hailuo 2.3 Fast, and Hailuo 02 are legacy V1 models. The currently documented speech and image IDs include speech-2.8-hd / speech-2.8-turbo and image-01. For music, the Music Generation and Lyrics Generation schemas remain published and still document music-3.0, music-2.6, and music-cover. That documents the endpoint contract, not account entitlement. MiniMax’s August 20, 2026 access notice says the paid Music Generation and Lyrics Generation APIs are unavailable to new users, existing paying users can continue to use the current services, and music-3.0-free, music-2.6-free, and music-cover-free are discontinued. Verify the actual key and account in the official console before building or purchasing around these APIs.
MiniMax API Pricing Snapshot
The following pay-as-you-go language rates were verified against MiniMax’s official pricing page on July 31, 2026. Rates are per one million tokens and can change; check the official page before budgeting.
| Model and tier | Input | Output | Cache read |
|---|---|---|---|
| M3 Standard, context ≤512K | $0.30 | $1.20 | $0.06 |
| M3 Standard, context >512K | $0.60 | $2.40 | $0.12 |
| M3 Priority, context ≤512K | $0.45 | $1.80 | $0.09 |
| M3 Priority, context >512K | $0.90 | $3.60 | $0.18 |
| M2.7 | $0.30 | $1.20 | $0.06 |
| M2.7-highspeed | $0.60 | $2.40 | $0.06 |
Priority routing is documented at 1.5 times the Standard rate and is selected with service_tier: "priority". It changes routing and price; it is not a promise that every request will have the same latency.
Worked cost example
At the M3 Standard ≤512K rates, a request using 20,000 input tokens and 2,000 output tokens costs approximately $0.0084: (20,000 × $0.30 / 1,000,000) + (2,000 × $1.20 / 1,000,000). Cache operations, Server Tools, retries, and other media endpoints are separate.
Current media examples include MiniMax-H3 at $0.08 per generated second for 768P or $0.13 for 2K, image-01 at $0.0035 per image, speech-2.8-turbo at $60 per million characters, speech-2.8-hd at $100 per million characters, and a published Music 3.0 rate of $0.15 for a clip up to five minutes for accounts that remain eligible. A listed rate does not establish access for a new account. Use our maintained MiniMax API pricing guide and the official pay-as-you-go page for the latest billing terms.
Rate Limits and Common MiniMax API Errors
| Model family | Requests per minute | Tokens per minute |
|---|---|---|
MiniMax-M3 | 200 RPM | 10,000,000 TPM |
MiniMax-M2.7 and MiniMax-M2.7-highspeed | 500 RPM | 20,000,000 TPM |
| Listed M2.5, M2.1, and M2 legacy models | 500 RPM | 20,000,000 TPM |
Video limits use a separate table: Video Generation V2 for MiniMax-H3 is capped by maximum concurrent tasks—2 on the free tier and 15 on the paid tier—rather than the legacy Hailuo RPM rule.
These are published platform limits, not a target concurrency setting. Add bounded exponential backoff with jitter for retryable failures, respect any server guidance, and cap both attempts and elapsed time. Do not automatically retry authentication, invalid-parameter, balance, or safety errors.
| Code | Meaning | First action |
|---|---|---|
1001 | Timeout | Retry only when the operation is safe and your total deadline allows it. |
1002 | Rate limit | Slow down, queue work, and apply bounded backoff. |
1004 / 2049 | Unauthorized / invalid API key | Check the key, regional host, header, and credential type. |
1008 | Insufficient balance | Check the wallet or billing resource. |
1039 | Token limit | Reduce input or output and recalculate the context budget. |
2013 | Invalid parameters | Validate the request against the exact endpoint schema. |
2056 | Token Plan usage window exceeded | Review the subscription quota and billing window. |
Use the dedicated MiniMax API rate-limits guide and MiniMax API error-code guide for retry decisions and the complete documented code list.
Security and Production Checklist
- Keep credentials on a trusted backend and rotate them after suspected exposure.
- Set explicit connect, read, and total request timeouts.
- Validate both transport-level and application-level success.
- Log request IDs, model IDs, latency, token usage, retry count, and sanitized error codes—never raw keys.
- Apply rate limiting, budgets, and per-user quotas before calling MiniMax.
- Validate tool arguments in your application; never execute model-generated commands without authorization and policy checks.
- Minimize prompts and uploads. Do not send secrets, private identifiers, regulated data, or proprietary files without an approved data-handling basis.
- Pin and test SDK versions, then regression-test the fields your application depends on.
- Use idempotency controls or your own task ledger for asynchronous and retryable workflows.
- Monitor price, model, endpoint, and deprecation changes independently of application releases.
Continue with the production MiniMax API application guide, the site’s security hub, and what not to paste into MiniMax AI.
How We Verified This Guide
Core verification date: August 13, 2026. Music API availability rechecked: August 20, 2026. We compared the model catalog, quick-start preparation, OpenAI-compatible Chat Completions, Responses, Anthropic Messages, H3 Video Generation V2, Music and Lyrics access notices, rate limits, error codes, and pay-as-you-go pricing against MiniMax’s official public documentation. The code examples were schema-reviewed for the documented interfaces. No private API key was used and the examples were not live-called from this editorial environment.
- Official prerequisites and key setup
- Official SDK quick start
- Official model catalog
- OpenAI-compatible Chat Completions reference
- Responses API reference
- Anthropic Messages reference
- Official rate limits
- Official error codes
- Official pay-as-you-go pricing
- Official Music Generation reference and August 20 access notice
- Official Lyrics Generation reference and August 20 access notice
- Official Video Generation V2 create reference
- Official Video Generation V2 query reference
Independent MiniMax M3 evidence
Our separate frozen-dataset tests cover short deterministic API endpoint parity, tool calling, Standard vs Priority latency, and M3 vision. Each test has its own stated methodology and limitations; none proves universal performance or full semantic parity.
Detailed MiniMax API Guides
| Topic | Guides |
|---|---|
| Keys and billing | API Key and base URL · Token Plan and Subscription Key |
| Text interfaces | OpenAI compatibility · Anthropic compatibility · Responses API |
| Tools and efficiency | Function calling · Prompt caching · Server Tools web search |
| Multimodal and files | M3 multimodal API · File management · Image API |
| Speech and voice | TTS pronunciation and subtitles · Voice Design API · WebSocket TTS |
| Music | Music API · Lyrics API · Music Cover API |
MiniMax API FAQ
What is the MiniMax API?
It is MiniMax’s developer platform for calling hosted language and media services from software. Text integrations can use OpenAI-compatible, Responses-compatible, or Anthropic-compatible request styles; media services use separate endpoints.
Where can I get a MiniMax API key?
Create it through MiniMax’s official Open Platform after registration. MiniMax-AI.chat is an independent guide and cannot create, display, recover, or manage your credential.
What is the MiniMax API base URL?
For international access, use https://api.minimax.io/v1 with OpenAI-compatible clients or https://api.minimax.io/anthropic with Anthropic-compatible clients. China-region hosts use api.minimaxi.com.
Is the MiniMax API OpenAI compatible?
MiniMax provides OpenAI-compatible Chat Completions, Responses, and model-listing routes. Compatibility does not imply that every OpenAI parameter or response field has identical behavior, so validate the exact fields you use.
Should I use OpenAI Chat or Anthropic Messages?
Use Chat Completions when you already rely on the OpenAI SDK or message format. Use Anthropic Messages when your stack uses the Anthropic SDK or its supported thinking and tool workflows. MiniMax currently presents the Anthropic SDK in its primary quick start.
Is the MiniMax API free?
Do not assume free production use. MiniMax publishes pay-as-you-go rates and separate Token Plans; account promotions and credits can change. Check the official billing page for your account and region.
What is the difference between an API Key and a Subscription Key?
A pay-as-you-go API Key spends from the relevant Open Platform wallet. A Subscription Key spends eligible Token Plan quota and then eligible Credits under MiniMax’s documented rules.
Which model should a new MiniMax API application use?
MiniMax-M3 is the current general-purpose model in the verified catalog and supports text, images, video, and tools on documented routes. Choose M2.7 or its high-speed variant only when its capability, speed, price, or compatibility better matches a tested workload.
Can I expose a MiniMax API key in frontend JavaScript?
No. A browser bundle is visible to users and attackers. Send requests through your own authenticated backend, keep the key in a secret manager, and enforce user-level quotas before forwarding calls.
Why can an HTTP 200 response still fail?
Some MiniMax operations expose application-level status information such as base_resp inside the response body. Validate both the transport status and the documented success fields for the exact endpoint.
Next step: create a server-side key, run one short request, record the model and usage fields, then add timeouts, quotas, validation, and monitoring before connecting the API to real users.
