MiniMax API Guide: Get a Key and Make Your First Call

Last verified: August 13, 2026. Music API availability rechecked: August 20, 2026.

The MiniMax API gives developers server-side access to MiniMax language, speech, video, image, music, voice, and file services. For a new text integration, use MiniMax-M3 with either the OpenAI-compatible base URL https://api.minimax.io/v1 or the Anthropic-compatible base URL https://api.minimax.io/anthropic.

Independent guide: MiniMax-AI.chat is not the official MiniMax website, does not issue API keys, and cannot access your account. This page is a dated, docs-verified reference. Confirm the schema for the exact endpoint you use before production deployment.

MiniMax API Quick Reference

What you needCurrent value
International OpenAI-compatible base URLhttps://api.minimax.io/v1
International Anthropic-compatible base URLhttps://api.minimax.io/anthropic
OpenAI Chat endpointPOST /v1/chat/completions
Responses endpointPOST /v1/responses
Anthropic Messages endpointPOST /anthropic/v1/messages
AuthenticationAuthorization: Bearer YOUR_API_KEY
Current general-purpose modelMiniMax-M3
China-region OpenAI / Anthropic baseshttps://api.minimaxi.com/v1 / https://api.minimaxi.com/anthropic

The base URL belongs in your SDK client configuration. Do not append /v1 twice. With the Anthropic SDK, set the base URL to https://api.minimax.io/anthropic; the SDK adds the Messages route.

What This MiniMax API Guide Covers

How to Get a MiniMax API Key

Create the credential in MiniMax’s official Open Platform after registration. The platform separates a pay-as-you-go API Key from the Subscription Key used by eligible Token Plan quota and Credits. Use the key that matches the billing resource you intend to spend.

  1. Open the official prerequisites guide and sign in to the correct regional platform.
  2. Create a pay-as-you-go API Key or the appropriate Subscription Key.
  3. Copy it once into a secret manager or a local environment variable.
  4. Never place the key in browser-side JavaScript, a mobile bundle, a public repository, a screenshot, or a support post.
  5. Use separate development and production credentials so usage and incident response remain attributable.
CredentialBilling sourceBest fit
Pay-as-you-go API KeyOpen Platform wallet for the account or teamMetered development and production requests
Subscription KeyEligible Token Plan quota, then eligible CreditsToken Plan seats and prepaid Credits

See the focused guides to MiniMax API keys and base URLs and API Key vs Subscription Key for setup and billing details.

Store the key safely

# macOS or Linux
export MINIMAX_API_KEY="replace-with-your-server-side-key"

# PowerShell
$env:MINIMAX_API_KEY="replace-with-your-server-side-key"

Do not commit a real .env file. Add it to .gitignore and use your hosting provider’s secret-management feature in production.

Make Your First MiniMax API Call in Five Minutes

The examples below use the international OpenAI-compatible route and MiniMax-M3. They intentionally request a short answer. Replace the prompt, output limit, and error handling for your application.

cURL: send a real Chat Completions request

curl https://api.minimax.io/v1/chat/completions \
  -H "Authorization: Bearer $MINIMAX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMax-M3",
    "messages": [
      {
        "role": "user",
        "content": "Reply with one short sentence confirming the API is connected."
      }
    ],
    "max_completion_tokens": 128,
    "thinking": {
      "type": "disabled"
    }
  }'

A successful response normally includes an assistant message plus token-usage data. Validate the HTTP status, the expected response object, and any MiniMax base_resp status fields returned by the specific endpoint. HTTP 200 alone should not be your only application-level success check.

Python with the OpenAI SDK

Install the client with pip install openai, then run:

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.minimax.io/v1",
    api_key=os.environ["MINIMAX_API_KEY"],
)

response = client.chat.completions.create(
    model="MiniMax-M3",
    messages=[
        {
            "role": "user",
            "content": "Reply with one short sentence confirming the API is connected.",
        }
    ],
    max_completion_tokens=128,
    extra_body={"thinking": {"type": "disabled"}},
)

print(response.choices[0].message.content)
print(response.usage)

Use max_completion_tokens for OpenAI-compatible Chat Completions. The older max_tokens parameter is deprecated on that interface.

Node.js with the OpenAI SDK

Install the client with npm install openai. Keep this code on your server, not in frontend JavaScript:

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.minimax.io/v1",
  apiKey: process.env.MINIMAX_API_KEY,
});

const response = await client.chat.completions.create({
  model: "MiniMax-M3",
  messages: [
    {
      role: "user",
      content: "Reply with one short sentence confirming the API is connected.",
    },
  ],
  max_completion_tokens: 128,
});

console.log(response.choices[0].message.content);
console.log(response.usage);

Python with the Anthropic SDK

MiniMax’s current quick-start documentation recommends the Anthropic SDK route. Install it with pip install anthropic:

import os
import anthropic

client = anthropic.Anthropic(
    base_url="https://api.minimax.io/anthropic",
    api_key=os.environ["MINIMAX_API_KEY"],
)

message = client.messages.create(
    model="MiniMax-M3",
    max_tokens=128,
    messages=[
        {
            "role": "user",
            "content": "Reply with one short sentence confirming the API is connected.",
        }
    ],
)

for block in message.content:
    if block.type == "text":
        print(block.text)

Do not mechanically rename this parameter: max_tokens remains correct for Anthropic Messages. The three current output-limit names are max_completion_tokens for OpenAI Chat, max_output_tokens for Responses, and max_tokens for Anthropic Messages.

Which MiniMax API Interface Should You Use?

InterfaceRouteOutput-limit fieldChoose it when
OpenAI Chat CompletionsPOST /v1/chat/completionsmax_completion_tokensYou already use the OpenAI SDK or Chat Completions message format.
OpenAI ResponsesPOST /v1/responsesmax_output_tokensYour application is built around the Responses object and its event model.
Anthropic MessagesPOST /anthropic/v1/messagesmax_tokensYou use the Anthropic SDK, Messages format, or supported thinking and tool workflows.

Compatibility does not mean every upstream parameter or response field behaves identically. Validate the fields your application uses against the exact MiniMax operation reference. Start with our OpenAI-compatible API guide, Anthropic-compatible API guide, or Responses API guide.

Legacy migration note: POST /v1/text/chatcompletion_v2 is still documented but marked Deprecated. Do not use it as the starting point for a new integration; migrate maintained text applications to Chat Completions, Responses, or Anthropic Messages.

Current MiniMax API Models

The official model catalog listed the following language models when this page was verified. Context figures combine input and output; they do not guarantee that every client, plan, or request will expose the maximum.

Model IDPublished contextInputsStatus and use
MiniMax-M31,000,000 tokensText, images, video, toolsCurrent general-purpose model for coding, agents, tools, and long-context work.
MiniMax-M2.7204,800 tokensText and toolsActive text model for engineering and professional workflows.
MiniMax-M2.7-highspeed204,800 tokensText and toolsSpeed-focused hosted M2.7 option.
MiniMax-M2.5 variants204,800 tokensText and toolsLegacy; preserve only where an existing integration requires them.
MiniMax-M2.1 variants and MiniMax-M2204,800 tokensText and toolsLegacy.

For MiniMax-M3, MiniMax documents a recommended output limit of 131,072 tokens and a hard maximum of 524,288. For other listed language models, the documented recommendation is 65,536 with a 204,800 maximum. Set a smaller explicit limit unless a tested workload needs more, and leave enough room for the input inside the total context budget.

See the detailed MiniMax M3 guide and MiniMax M2.7 guide. Hosted API availability and downloadable model-weight licensing are separate questions.

Parameters, Streaming, Tools, and Multimodal Input

  • Streaming: stream: true is documented for Chat Completions, Responses, and Messages. Your client must process the event format for the selected interface.
  • Tools: use the modern tools field. The older function_call field is not supported. tool_choice supports auto and none.
  • One output per request: the Chat Completions n field supports 1.
  • Tool history: in a multi-turn tool workflow, retain and resend the complete assistant response, including tool calls and supported thinking fields.
  • Multimodal: MiniMax-M3 accepts text, image, and video input on documented routes. Audio input is not supported by Chat Completions. M2.x models are text-and-tools models.
  • Thinking: thinking behavior and defaults differ by interface. Configure it deliberately and preserve the returned state when the workflow requires it.

For implementation details, use the dedicated guides to MiniMax function calling, prompt caching, M3 multimodal input, and Anthropic Server Tools web search. Server Tools is currently a beta Anthropic Messages feature; the documented web_search tool costs $0.01 per request.

MiniMax API Endpoint Map

All routes below use https://api.minimax.io as the international host. A chat endpoint does not generate every media type: text, speech, video, image, music, voice, and managed files use separate schemas.

Language endpoints

CapabilityMethod and route
List OpenAI-compatible modelsGET /v1/models
Chat CompletionsPOST /v1/chat/completions
ResponsesPOST /v1/responses
Count Responses input tokensPOST /v1/responses/input_tokens
Anthropic MessagesPOST /anthropic/v1/messages
Count Anthropic input tokensPOST /anthropic/v1/messages/count_tokens

Media, voice, and file endpoints

CapabilityMethod and routeWorkflow
Synchronous text to speechPOST /v1/t2a_v2Return speech over HTTP; a WebSocket interface is also documented.
Asynchronous long speechPOST /v1/t2a_async_v2Create a task, then poll /v1/query/t2a_async_query_v2.
Video Generation V2 (MiniMax-H3)POST /v2/video_generationCreate a task, then query GET /v2/query/video_generation/{task_id}; Pay as You Go is required.
Legacy Hailuo video generationPOST /v1/video_generationCreate a task, then poll GET /v1/query/video_generation?task_id=....
Image generationPOST /v1/image_generationText-to-image or supported subject-reference workflow.
Music generationPOST /v1/music_generationSong, instrumental, or supported cover workflow.
Lyrics generationPOST /v1/lyrics_generationCreate, edit, or continue lyrics.
Voice cloningPOST /v1/voice_cloneCreate a voice from eligible uploaded audio.
Voice designPOST /v1/voice_designDesign a voice from a description.
Upload a managed filePOST /v1/files/uploadUpload with the purpose required by the destination endpoint.
Retrieve file metadataGET /v1/files/retrieve?file_id=...Read metadata or a generated-file URL.
Retrieve file contentGET /v1/files/retrieve_content?file_id=...Download managed content where supported.

The model catalog verified on August 13, 2026 lists MiniMax-H3 as the current video model; it uses the V2 API. Hailuo 2.3, Hailuo 2.3 Fast, and Hailuo 02 are legacy V1 models. The currently documented speech and image IDs include speech-2.8-hd / speech-2.8-turbo and image-01. For music, the Music Generation and Lyrics Generation schemas remain published and still document music-3.0, music-2.6, and music-cover. That documents the endpoint contract, not account entitlement. MiniMax’s August 20, 2026 access notice says the paid Music Generation and Lyrics Generation APIs are unavailable to new users, existing paying users can continue to use the current services, and music-3.0-free, music-2.6-free, and music-cover-free are discontinued. Verify the actual key and account in the official console before building or purchasing around these APIs.

MiniMax API Pricing Snapshot

The following pay-as-you-go language rates were verified against MiniMax’s official pricing page on July 31, 2026. Rates are per one million tokens and can change; check the official page before budgeting.

Model and tierInputOutputCache read
M3 Standard, context ≤512K$0.30$1.20$0.06
M3 Standard, context >512K$0.60$2.40$0.12
M3 Priority, context ≤512K$0.45$1.80$0.09
M3 Priority, context >512K$0.90$3.60$0.18
M2.7$0.30$1.20$0.06
M2.7-highspeed$0.60$2.40$0.06

Priority routing is documented at 1.5 times the Standard rate and is selected with service_tier: "priority". It changes routing and price; it is not a promise that every request will have the same latency.

Worked cost example

At the M3 Standard ≤512K rates, a request using 20,000 input tokens and 2,000 output tokens costs approximately $0.0084: (20,000 × $0.30 / 1,000,000) + (2,000 × $1.20 / 1,000,000). Cache operations, Server Tools, retries, and other media endpoints are separate.

Current media examples include MiniMax-H3 at $0.08 per generated second for 768P or $0.13 for 2K, image-01 at $0.0035 per image, speech-2.8-turbo at $60 per million characters, speech-2.8-hd at $100 per million characters, and a published Music 3.0 rate of $0.15 for a clip up to five minutes for accounts that remain eligible. A listed rate does not establish access for a new account. Use our maintained MiniMax API pricing guide and the official pay-as-you-go page for the latest billing terms.

Rate Limits and Common MiniMax API Errors

Model familyRequests per minuteTokens per minute
MiniMax-M3200 RPM10,000,000 TPM
MiniMax-M2.7 and MiniMax-M2.7-highspeed500 RPM20,000,000 TPM
Listed M2.5, M2.1, and M2 legacy models500 RPM20,000,000 TPM

Video limits use a separate table: Video Generation V2 for MiniMax-H3 is capped by maximum concurrent tasks—2 on the free tier and 15 on the paid tier—rather than the legacy Hailuo RPM rule.

These are published platform limits, not a target concurrency setting. Add bounded exponential backoff with jitter for retryable failures, respect any server guidance, and cap both attempts and elapsed time. Do not automatically retry authentication, invalid-parameter, balance, or safety errors.

CodeMeaningFirst action
1001TimeoutRetry only when the operation is safe and your total deadline allows it.
1002Rate limitSlow down, queue work, and apply bounded backoff.
1004 / 2049Unauthorized / invalid API keyCheck the key, regional host, header, and credential type.
1008Insufficient balanceCheck the wallet or billing resource.
1039Token limitReduce input or output and recalculate the context budget.
2013Invalid parametersValidate the request against the exact endpoint schema.
2056Token Plan usage window exceededReview the subscription quota and billing window.

Use the dedicated MiniMax API rate-limits guide and MiniMax API error-code guide for retry decisions and the complete documented code list.

Security and Production Checklist

  • Keep credentials on a trusted backend and rotate them after suspected exposure.
  • Set explicit connect, read, and total request timeouts.
  • Validate both transport-level and application-level success.
  • Log request IDs, model IDs, latency, token usage, retry count, and sanitized error codes—never raw keys.
  • Apply rate limiting, budgets, and per-user quotas before calling MiniMax.
  • Validate tool arguments in your application; never execute model-generated commands without authorization and policy checks.
  • Minimize prompts and uploads. Do not send secrets, private identifiers, regulated data, or proprietary files without an approved data-handling basis.
  • Pin and test SDK versions, then regression-test the fields your application depends on.
  • Use idempotency controls or your own task ledger for asynchronous and retryable workflows.
  • Monitor price, model, endpoint, and deprecation changes independently of application releases.

Continue with the production MiniMax API application guide, the site’s security hub, and what not to paste into MiniMax AI.

How We Verified This Guide

Core verification date: August 13, 2026. Music API availability rechecked: August 20, 2026. We compared the model catalog, quick-start preparation, OpenAI-compatible Chat Completions, Responses, Anthropic Messages, H3 Video Generation V2, Music and Lyrics access notices, rate limits, error codes, and pay-as-you-go pricing against MiniMax’s official public documentation. The code examples were schema-reviewed for the documented interfaces. No private API key was used and the examples were not live-called from this editorial environment.

Independent MiniMax M3 evidence

Our separate frozen-dataset tests cover short deterministic API endpoint parity, tool calling, Standard vs Priority latency, and M3 vision. Each test has its own stated methodology and limitations; none proves universal performance or full semantic parity.

Detailed MiniMax API Guides

TopicGuides
Keys and billingAPI Key and base URL · Token Plan and Subscription Key
Text interfacesOpenAI compatibility · Anthropic compatibility · Responses API
Tools and efficiencyFunction calling · Prompt caching · Server Tools web search
Multimodal and filesM3 multimodal API · File management · Image API
Speech and voiceTTS pronunciation and subtitles · Voice Design API · WebSocket TTS
MusicMusic API · Lyrics API · Music Cover API

MiniMax API FAQ

What is the MiniMax API?

It is MiniMax’s developer platform for calling hosted language and media services from software. Text integrations can use OpenAI-compatible, Responses-compatible, or Anthropic-compatible request styles; media services use separate endpoints.

Where can I get a MiniMax API key?

Create it through MiniMax’s official Open Platform after registration. MiniMax-AI.chat is an independent guide and cannot create, display, recover, or manage your credential.

What is the MiniMax API base URL?

For international access, use https://api.minimax.io/v1 with OpenAI-compatible clients or https://api.minimax.io/anthropic with Anthropic-compatible clients. China-region hosts use api.minimaxi.com.

Is the MiniMax API OpenAI compatible?

MiniMax provides OpenAI-compatible Chat Completions, Responses, and model-listing routes. Compatibility does not imply that every OpenAI parameter or response field has identical behavior, so validate the exact fields you use.

Should I use OpenAI Chat or Anthropic Messages?

Use Chat Completions when you already rely on the OpenAI SDK or message format. Use Anthropic Messages when your stack uses the Anthropic SDK or its supported thinking and tool workflows. MiniMax currently presents the Anthropic SDK in its primary quick start.

Is the MiniMax API free?

Do not assume free production use. MiniMax publishes pay-as-you-go rates and separate Token Plans; account promotions and credits can change. Check the official billing page for your account and region.

What is the difference between an API Key and a Subscription Key?

A pay-as-you-go API Key spends from the relevant Open Platform wallet. A Subscription Key spends eligible Token Plan quota and then eligible Credits under MiniMax’s documented rules.

Which model should a new MiniMax API application use?

MiniMax-M3 is the current general-purpose model in the verified catalog and supports text, images, video, and tools on documented routes. Choose M2.7 or its high-speed variant only when its capability, speed, price, or compatibility better matches a tested workload.

Can I expose a MiniMax API key in frontend JavaScript?

No. A browser bundle is visible to users and attackers. Send requests through your own authenticated backend, keep the key in a secret manager, and enforce user-level quotas before forwarding calls.

Why can an HTTP 200 response still fail?

Some MiniMax operations expose application-level status information such as base_resp inside the response body. Validate both the transport status and the documented success fields for the exact endpoint.

Next step: create a server-side key, run one short request, record the model and usage fields, then add timeouts, quotas, validation, and monitoring before connecting the API to real users.