Last verified: July 16, 2026.
Independent tutorial notice: MiniMax-AI.chat is not an official MiniMax website. The examples below use MiniMax’s documented API interfaces, but they are independent implementation examples rather than MiniMax support material. Test every request in your own account, review first-party documentation before deployment, and never paste a real secret key into this page or public client code.
This tutorial shows how to build a server-side AI application with the MiniMax API. Node.js is the primary implementation. A smaller Python example follows for teams using Python services. The guide covers architecture, a non-streaming route, Server-Sent Events streaming, input validation, secrets, retry policy, media task workflows, evaluation, observability, cost controls, and deployment checks.
For model and endpoint lookup rather than implementation, use the separate MiniMax API reference hub. For MiniMax Code as an end-user coding product, use the MiniMax Code guide. This page focuses on applications you operate through the API.
What you will build
The working example is a small Express service with two routes:
POST /api/chatvalidates a conversation and returns one completed answer.POST /api/chat/streamvalidates the same input and streams user-visible answer text over Server-Sent Events.
The service calls MiniMax-M3 through MiniMax’s OpenAI-compatible endpoint. The same production boundaries—server-side authentication, user authorization, validation, quotas, monitoring, and evaluation—also apply if you choose the Anthropic-compatible interface.
Recommended application architecture
- Browser or mobile client: collects user input and displays progress, but never receives the MiniMax secret.
- Your application backend: authenticates the user, validates input, enforces quotas, chooses the model, and calls MiniMax.
- Policy and tool layer: decides which requests are allowed and which tool actions require authorization or confirmation.
- MiniMax API: returns a language response or creates a modality-specific generation task.
- Storage and observability: retain only the operational data you have deliberately approved, with access controls and a defined deletion schedule.
Do not call a paid provider directly from public JavaScript. Domain restrictions cannot turn a browser-visible API key into a secret. A backend also gives you one place to apply tenant isolation, spend limits, abuse controls, request tracing, and model migrations.
Step 1: Create the correct MiniMax credential
Register or sign in through MiniMax’s official platform, create the credential required by your billing route, and add eligible resources. MiniMax distinguishes a pay-as-you-go API Key from a Subscription Key used with Token Plan subscriptions and purchased Credits. Follow the official prerequisites instead of assuming those credentials are interchangeable.
- Use a separate secret for development, staging, and production where the account permits it.
- Store production secrets in the deployment platform’s secret manager.
- Restrict who can read or rotate each secret.
- Never commit
.env; add it to.gitignore. - Rotate a secret immediately if it appears in a repository, log, screenshot, ticket, or client bundle.
Step 2: Set up the Node.js project
This example uses JavaScript modules, Node.js 20 or later, Express, the OpenAI SDK, Zod validation, Helmet security headers, an application rate limiter, and dotenv for local development.
mkdir minimax-app
cd minimax-app
npm init -y
npm install openai express zod helmet express-rate-limit dotenv
Add the module type and scripts without replacing the dependencies that npm added to package.json:
npm pkg set type=module
npm pkg set scripts.start="node src/server.js"
npm pkg set scripts.dev="node --watch src/server.js"
Create a local .env file:
MINIMAX_API_KEY=replace_with_your_secret
MINIMAX_MODEL=MiniMax-M3
PORT=3000
MiniMax’s verified OpenAI-compatible base URL is https://api.minimax.io/v1. The official OpenAI SDK page documents that base and the MiniMax-M3 model ID.
Step 3: Create one server-side MiniMax client
Create src/minimax.js:
import OpenAI from "openai";
const apiKey = process.env.MINIMAX_API_KEY;
if (!apiKey) {
throw new Error("MINIMAX_API_KEY is required");
}
export const model = process.env.MINIMAX_MODEL || "MiniMax-M3";
export const minimax = new OpenAI({
apiKey,
baseURL: "https://api.minimax.io/v1",
timeout: 45_000,
maxRetries: 0
});
The SDK’s automatic retries are disabled here because the application will own a visible and testable retry policy. Do not combine several SDK, proxy, queue, and load-balancer retry layers without a shared attempt budget.
Step 4: Validate input before calling the model
Create src/schemas.js. These limits are application choices, not MiniMax platform limits. They keep the demonstration bounded and should be adjusted from measured product needs.
import { z } from "zod";
const messageSchema = z.object({
role: z.enum(["user", "assistant"]),
content: z.string().trim().min(1).max(20_000)
});
export const chatSchema = z.object({
messages: z.array(messageSchema).min(1).max(30)
});
Do not accept a client-supplied system role, model ID, output ceiling, service tier, or tool definition unless your product explicitly authorizes each field. Keep high-impact configuration on the server.
Step 5: Add a non-streaming chat route
Create src/server.js:
import "dotenv/config";
import crypto from "node:crypto";
import express from "express";
import helmet from "helmet";
import rateLimit from "express-rate-limit";
import { minimax, model } from "./minimax.js";
import { chatSchema } from "./schemas.js";
const app = express();
app.use(helmet());
app.use(express.json({ limit: "64kb" }));
app.use(rateLimit({
windowMs: 60_000,
limit: 20,
standardHeaders: "draft-7",
legacyHeaders: false
}));
app.post("/api/chat", async (req, res) => {
const requestId = crypto.randomUUID();
const parsed = chatSchema.safeParse(req.body);
if (!parsed.success) {
return res.status(400).json({
error: "invalid_request",
requestId
});
}
const startedAt = Date.now();
try {
const completion = await minimax.chat.completions.create({
model,
messages: [
{
role: "system",
content: "Answer clearly. State uncertainty and do not invent sources."
},
...parsed.data.messages
],
thinking: { type: "disabled" },
max_completion_tokens: 800,
temperature: 0.4
});
const answer = completion.choices[0]?.message?.content;
if (!answer) {
throw new Error("MiniMax returned no user-visible answer");
}
console.info(JSON.stringify({
event: "minimax_completion",
requestId,
model: completion.model,
finishReason: completion.choices[0]?.finish_reason,
promptTokens: completion.usage?.prompt_tokens,
completionTokens: completion.usage?.completion_tokens,
durationMs: Date.now() - startedAt
}));
return res.json({
requestId,
model: completion.model,
answer
});
} catch (error) {
console.error(JSON.stringify({
event: "minimax_error",
requestId,
status: error?.status,
durationMs: Date.now() - startedAt
}));
return res.status(502).json({
error: "model_provider_error",
requestId
});
}
});
const port = Number(process.env.PORT || 3000);
app.listen(port, () => {
console.log(`Server listening on port ${port}`);
});
The example explicitly disables M3 thinking so the returned content is suitable for a basic user-facing chat. MiniMax documents that M2.x thinking cannot be disabled, even if the disabled value is sent. If you change the model, inspect its response format and test what your interface exposes.
Authentication for your own users is omitted only to keep the code readable. Add a real session, token, or service-authentication layer before the rate limiter and chat routes. IP rate limiting alone is not tenant authorization and can be unreliable behind proxies unless proxy trust is configured carefully.
Step 6: Stream answers with Server-Sent Events
Streaming improves perceived responsiveness because the client can render text while generation continues. It does not guarantee lower total completion time. Add this route before app.listen:
app.post("/api/chat/stream", async (req, res) => {
const requestId = crypto.randomUUID();
const parsed = chatSchema.safeParse(req.body);
if (!parsed.success) {
return res.status(400).json({ error: "invalid_request", requestId });
}
const controller = new AbortController();
res.on("close", () => controller.abort());
res.status(200);
res.setHeader("Content-Type", "text/event-stream; charset=utf-8");
res.setHeader("Cache-Control", "no-cache, no-transform");
res.setHeader("Connection", "keep-alive");
res.setHeader("X-Accel-Buffering", "no");
res.flushHeaders();
try {
const stream = await minimax.chat.completions.create(
{
model,
messages: [
{ role: "system", content: "Answer clearly and state uncertainty." },
...parsed.data.messages
],
thinking: { type: "disabled" },
max_completion_tokens: 800,
temperature: 0.4,
stream: true,
stream_options: { include_usage: true }
},
{ signal: controller.signal }
);
for await (const chunk of stream) {
const delta = chunk.choices[0]?.delta?.content;
if (delta) {
res.write(`data: ${JSON.stringify({ type: "text", delta })}\n\n`);
}
if (chunk.usage) {
res.write(`data: ${JSON.stringify({
type: "usage",
usage: chunk.usage
})}\n\n`);
}
}
res.write(`data: ${JSON.stringify({ type: "done", requestId })}\n\n`);
res.end();
} catch (error) {
if (!res.writableEnded) {
res.write(`data: ${JSON.stringify({
type: "error",
requestId
})}\n\n`);
res.end();
}
}
});
Some reverse proxies buffer responses unless streaming is configured explicitly. Test through the same CDN, load balancer, serverless runtime, and browser path used in production. Stop provider work when the downstream connection closes, and set an application deadline so abandoned requests do not consume resources indefinitely.
Step 7: Add a bounded retry policy
Retry transient failures, not every failure. MiniMax’s error reference identifies retry-oriented provider codes such as 1000, 1001, 1002, 1024, and 1033. Authentication failures, invalid parameters, and insufficient balance require correction rather than repeated traffic.
const retryableStatuses = new Set([408, 429, 500, 502, 503, 504]);
const retryableProviderCodes = new Set([1000, 1001, 1002, 1024, 1033]);
function providerCode(error) {
return Number(
error?.error?.base_resp?.status_code ??
error?.body?.base_resp?.status_code ??
NaN
);
}
function delay(ms) {
return new Promise((resolve) => setTimeout(resolve, ms));
}
export async function withRetry(operation, maxAttempts = 3) {
let lastError;
for (let attempt = 1; attempt <= maxAttempts; attempt += 1) {
try {
return await operation();
} catch (error) {
lastError = error;
const retryable =
retryableStatuses.has(error?.status) ||
retryableProviderCodes.has(providerCode(error));
if (!retryable || attempt === maxAttempts) {
throw error;
}
const exponentialMs = 500 * 2 ** (attempt - 1);
const jitterMs = Math.floor(Math.random() * 250);
await delay(exponentialMs + jitterMs);
}
}
throw lastError;
}
Save the helper as src/retry.js, import it into server.js, and wrap the non-streaming provider call:
import { withRetry } from "./retry.js";
const completion = await withRetry(() =>
minimax.chat.completions.create({
model,
messages,
thinking: { type: "disabled" },
max_completion_tokens: 800,
temperature: 0.4
})
);
In this short fragment, messages represents the validated server-built array from the earlier route. Do not restart an SSE response automatically after text has reached the client; that can duplicate output. Retry only before streaming begins, or expose a deliberate user retry action.
Use one total deadline across attempts. Add the attempt number to logs, respect Retry-After when supplied, and cap concurrency before rate limits are reached. The official rate-limit tables vary by model and modality, so do not use one copied RPM value for every route.
Do not blindly retry an ambiguous media-task creation timeout. The first request may have created a chargeable task even when your application missed the response. Record an application operation ID, reconcile the outcome where the provider workflow permits it, and prevent duplicate clicks at the product layer.
Python alternative
Teams using Python can call the same OpenAI-compatible interface. Keep this code on the server:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["MINIMAX_API_KEY"],
base_url="https://api.minimax.io/v1",
timeout=45.0,
max_retries=0,
)
response = client.chat.completions.create(
model="MiniMax-M3",
messages=[
{"role": "system", "content": "Answer clearly and state uncertainty."},
{"role": "user", "content": "Give me three test cases for a login form."},
],
max_completion_tokens=800,
temperature=0.4,
extra_body={"thinking": {"type": "disabled"}},
)
print(response.choices[0].message.content)
The Python snippet is intentionally small. Apply the same validation, authentication, quotas, timeouts, retries, structured logs, and evaluations used in the Node service.
Choosing a model without coupling the product to it
MiniMax’s model catalog lists MiniMax-M3, MiniMax-M2.7, and MiniMax-M2.7-highspeed as primary language-model choices, while several earlier M2-family IDs are marked legacy. Check the official model catalog before changing configuration.
| Application need | Evaluation starting point | Do not assume |
|---|---|---|
| Coding, agentic workflows, long context, image/video understanding | Evaluate MiniMax-M3 | A one-million-token context makes every large prompt accurate, fast, or economical. |
| Text and tool workflows | Evaluate MiniMax-M2.7 | Published benchmark strength guarantees performance on your repository or domain. |
| Lower-latency M2.7 option | Compare MiniMax-M2.7-highspeed against M2.7 | Speed removes the need to test quality, cost, concurrency, and tail latency. |
Read the model ID from environment configuration, return the provider-reported model ID with each response, and record it in traces. Promote a model only after it passes the same evaluation set as the version it replaces.
Tool calling and agent actions
A model-generated tool call is a proposed action, not authorization. Your code must decide whether the user may perform it and whether the arguments are valid.
- Expose a small allowlist of tools with narrow descriptions and strict JSON schemas.
- Parse and validate every argument independently of the model.
- Bind the operation to the authenticated user and tenant.
- Require explicit confirmation for purchases, deletion, publication, messages, permissions, and other high-impact actions.
- Apply timeouts and least-privilege credentials to each tool.
- Return a compact, untrusted tool result to the model; do not treat retrieved content as instructions.
- Limit turns, tool calls, tokens, wall-clock duration, and spend for every run.
MiniMax instructs M3 developers to preserve the complete assistant response across multi-turn tool conversations, including the fields needed for the reasoning chain. Keep that state inside the orchestration layer and do not expose internal thinking fields to end users or analytics by default. Follow the first-party Tool Use and Interleaved Thinking guide for protocol-specific message handling.
For a broader protocol explanation, see our MiniMax MCP guide. For code-focused use cases and boundaries, see MiniMax AI for code generation.
Adding retrieval without turning a long context window into a database
A large context window does not replace retrieval design. For document assistants:
- Authorize access before retrieval, not after generation.
- Parse and chunk documents with stable source IDs.
- Retrieve a bounded set of relevant passages.
- Provide citations and instruct the model to distinguish evidence from inference.
- Refuse or qualify an answer when the available evidence is insufficient.
- Evaluate retrieval recall separately from answer quality.
Our MiniMax RAG guide covers that architecture in greater depth. Keep this API tutorial focused on provider integration rather than duplicating the retrieval guide.
Image, video, speech, and music applications
Do not route every capability through Chat Completions. MiniMax has separate APIs for generated media:
| Workflow | Implementation pattern | Production concern |
|---|---|---|
| Synchronous speech | Send a bounded text-to-audio request and stream or return audio according to the T2A schema. | Validate text length, voice access, format, consent, and output storage. |
| Long-form speech | Create an asynchronous task, poll status with backoff, then retrieve the output file. | Persist task state and handle expiry, cancellation, and failed jobs. |
| Video | Create a video-generation task, poll by task_id, then retrieve the resulting file_id. | Avoid aggressive polling and duplicate task creation. |
| Image | Call the image-generation route with the operation’s text or reference-image schema. | Validate uploads, rights, size, format, and retention. |
| Music | Use the music-generation operation with an eligible generation or cover model. | Address lyrics, reference-audio rights, output rights, and product moderation. |
For video, asynchronous speech, and other jobs, store a record such as user_id, operation_id, provider_task_id, status, timestamps, attempt count, and output reference. Do not store the original prompt or media by default merely because a task table exists.
Security and privacy controls
Secrets and network boundaries
- Use server-side secrets and TLS.
- Separate environments and rotate credentials.
- Apply outbound allowlists where appropriate.
- Never return provider errors containing headers, credentials, account details, or raw upstream bodies to the client.
User and tenant isolation
- Authorize every conversation, file, task, and tool action.
- Use server-derived tenant IDs; do not trust a tenant ID submitted by the client.
- Enforce quotas per user and tenant in addition to IP controls.
- Prevent one tenant from retrieving another tenant’s task or generated file.
Data minimization
- Tell users when content is sent to an external AI provider.
- Remove secrets and unnecessary identifiers before a request.
- Define retention separately for prompts, outputs, uploads, traces, backups, support tickets, and analytics.
- Use redaction or metadata-only logging where content is not required.
- Review MiniMax’s API Privacy Policy alongside your own legal and security requirements.
An API integration is not automatically GDPR-, HIPAA-, SOC 2-, or sector-compliant. Compliance depends on the full system, parties, purpose, legal basis, contracts, locations, configuration, retention, access controls, monitoring, and operational practices. For user guidance, link to what not to paste into MiniMax AI.
Observability without logging everything
Capture enough metadata to diagnose reliability and cost:
- Application request ID and provider request ID when available.
- Environment, route, model requested, and model returned.
- Start time, time to first token for streams, total duration, and status.
- Prompt, completion, and cached-token usage when returned.
- Finish reason, retry count, provider error code, HTTP status, and cancellation.
- Tool name and result status without unrestricted argument or result bodies.
Do not log the Authorization header. Avoid raw prompts and responses unless a documented debugging or evaluation purpose requires them. If content logging is enabled, sample deliberately, redact sensitive fields, restrict access, define retention, and tell users where required.
Evaluation before deployment
Published benchmarks do not prove that a model meets your product requirements. Build a versioned test set from representative, permitted examples and measure the behavior that matters to your application.
| Dimension | Example measurement | Failure to include |
|---|---|---|
| Task quality | Rubric score, exact match, groundedness, code tests, or human preference | Optimizing only for fluency |
| Safety | Policy adherence, data leakage, prompt injection, and tool authorization tests | Testing only benign prompts |
| Reliability | Error rate, empty responses, truncation, retry rate, and task completion | Reporting one successful demo |
| Latency | P50 and P95 first-token and total latency under expected concurrency | Using one local request as a universal speed claim |
| Cost | Cost per successful task, including retries, output, media jobs, and storage | Comparing only headline input-token prices |
Run the same suite for every prompt, model, parameter, retrieval, or tool change. Record the exact model ID, date, dataset version, settings, region, and acceptance threshold. This guide does not publish a performance benchmark because no standardized test was run for this demonstration.
Cost and capacity planning
- Estimate input, cached input, and output separately for language workloads.
- Set output ceilings based on the task rather than the model’s maximum.
- Track cost per successful outcome, not only cost per request.
- Include retries, failed asynchronous jobs, polling, file storage, and egress.
- Use per-user and per-tenant budgets with alerts and hard stops.
- Load-test within approved quotas and request increased limits through MiniMax if justified.
Review MiniMax’s official pay-as-you-go pricing for billing units and our independent pricing guide for a cross-modality explanation. Recalculate before launch and after any model, prompt, output, traffic, or retry change.
Error-handling policy
| Failure | Retry? | Product behavior |
|---|---|---|
| Timeout, transient provider failure, or rate limit | Limited retries with backoff if the operation is safe to repeat | Show a stable request ID and a non-technical message. |
| Invalid or inactive key | No repeated retry | Alert the operator; never ask an end user for the provider secret. |
| Insufficient balance | No | Stop traffic or use an approved fallback after checking billing. |
| Invalid request | No | Fix server validation or the endpoint-specific mapping. |
| Sensitive input or output response | Do not attempt safeguard evasion | Apply product policy, offer safe reformulation where appropriate, or refuse. |
| Client disconnect | No automatic restart | Cancel upstream work where supported and record cancellation. |
Use the official MiniMax error-code definitions as the source of record. Our MiniMax API troubleshooting guide provides a broader diagnostic workflow.
Deployment checklist
- Provider secrets exist only in the deployment secret store.
- User authentication and tenant authorization run before provider calls.
- Request body, message count, content length, file type, and file size are validated.
- Per-user, per-tenant, and global spend and rate limits are active.
- Timeout, cancellation, retry, and circuit-breaker behavior have been tested.
- Streaming works through the production CDN and proxy path without unwanted buffering.
- Logs redact secrets and unnecessary personal or confidential content.
- Quality, safety, latency, reliability, and cost evaluations meet written thresholds.
- Model ID is configurable and provider-reported IDs are observable.
- Asynchronous tasks survive process restarts and reject cross-tenant access.
- Privacy notices, retention, deletion, incident response, and user support processes match the implemented data flow.
- Billing alerts and a controlled shutdown or fallback path have been rehearsed.
Frequently asked questions
Can I call the MiniMax API directly from a browser?
You should not expose a MiniMax secret in browser code. Use your own authenticated backend, enforce quotas there, and stream or return the permitted result to the client.
Should I use the OpenAI-compatible or Anthropic-compatible interface?
Use the protocol that fits your stack and required response semantics. MiniMax labels Anthropic SDK as recommended for language models and also documents an OpenAI-compatible SDK path. Do not assume every parameter behaves identically across providers or protocols.
Does MiniMax-M3 accept images and video?
MiniMax documents supported image and video input for M3 through its language interfaces. That capability analyzes supplied media; it is separate from the image- and video-generation APIs.
Can I treat a one-million-token context as a one-million-token prompt?
No. The published context window covers input and output, and usable capacity can depend on the interface and infrastructure. Large prompts also affect latency, cost, retrieval quality, and failure risk. Measure the real workload.
Should I retry every MiniMax error?
No. Retry only bounded transient failures. Invalid credentials, invalid parameters, insufficient balance, policy responses, and unsafe-to-repeat task creation require different handling.
Does using MiniMax make my app compliant?
No API alone establishes compliance. Assess the complete application, data flow, organizations, contracts, locations, purpose, user notices, access controls, logs, retention, security, and operational evidence.
Next implementation decisions
Once the server route is stable, choose one product-specific path instead of expanding a generic chatbot indefinitely:
- For grounded internal answers, build the authorization and retrieval layer described in the RAG guide.
- For coding workflows, define repository access, sandboxing, tests, and approval boundaries rather than granting an agent unrestricted execution.
- For speech, video, image, or music, implement the modality’s exact task schema and rights controls.
- For tool-driven agents, start with one low-risk read-only tool and add permissions only after evaluation.
The objective is not merely to receive a successful model response. A production MiniMax application must give users a reliable outcome while controlling secrets, data, permissions, cost, failures, and change.
