MiniMax WebSocket TTS is the synchronous speech route for applications that need audio chunks before an entire utterance has finished synthesizing. A server connects to wss://api.minimax.io/ws/v1/t2a_v2, authenticates with a MiniMax API key, starts a task, submits text, decodes the returned hexadecimal audio, and explicitly finishes the session.
Requires MiniMax Open Platform API. MiniMax-AI.chat is an independent resource and is not operated by MiniMax. The MiniMax-AI.chat demo is text-only and does not expose speech generation or a WebSocket connection.
Last verified: July 18, 2026. Endpoint names, event messages, model IDs, limits, pricing, and rate-limit statements on this page were checked against MiniMax’s official documentation. The examples were syntax-checked, but no paid synthesis request was submitted for this article.
MiniMax WebSocket TTS at a glance
| Item | Documented behavior | Implementation consequence |
|---|---|---|
| Transport | Secure WebSocket | Connect from a trusted backend, not public browser code. |
| Endpoint | wss://api.minimax.io/ws/v1/t2a_v2 | Send the API key in the Authorization: Bearer … handshake header. |
| Text limit | Up to 10,000 characters per synchronous request | Use the asynchronous route for substantially longer material. |
| Audio delivery | data.audio contains hex-encoded audio chunks | Decode with a hex decoder; do not treat the value as Base64. |
| Client events | task_start, one or more task_continue events, then task_finish | Respect the state machine instead of sending text immediately after opening the socket. |
| Server events | connected_success, task_started, continued results, task_finished, or task_failed | Gate every client action on the corresponding server acknowledgement. |
| Idle behavior | The connection closes if no new event is sent within 120 seconds after the last result | Finish promptly; do not keep an unused session open. |
| Speech rate limit | The general T2A table lists 60 requests per minute | Queue work and re-check the allowance attached to the production account. |
This page deliberately focuses on the WebSocket protocol and streaming client. For a broad API map, see the MiniMax API guide. For long scripts, task polling, downloadable result files, and sentence-level subtitle output, use the separate MiniMax Async TTS guide.
When to use WebSocket TTS
Choose the WebSocket route when an application benefits from receiving audio incrementally: a conversational agent, accessibility reader, interactive language exercise, or server-side narration service. The first playable bytes can be sent to a decoder while later chunks are arriving.
Do not choose it merely because the source text is long. The synchronous WebSocket guide caps a request at 10,000 characters. Async TTS accepts up to 50,000 characters in the direct text field or up to 1,000,000 characters through a supported uploaded file. Those modes solve different problems: WebSocket reduces delivery latency, while async TTS manages long-running jobs and downloadable artifacts.
| Requirement | Better route | Reason |
|---|---|---|
| Play speech while it is generated | WebSocket TTS | Audio arrives as incremental hex chunks. |
| Simple request/response integration | Synchronous HTTP TTS | No persistent bidirectional session is required. |
| Book, course, podcast script, or batch narration | Async TTS | Task polling, file input, result retrieval, and sentence-level subtitle output are documented. |
| Direct browser integration | Backend proxy required | Embedding a MiniMax API key in client-side JavaScript exposes the credential. |
Supported models and billing
The official WebSocket guide lists six model IDs: speech-2.8-hd, speech-2.8-turbo, speech-2.6-hd, speech-2.6-turbo, speech-02-hd, and speech-02-turbo. MiniMax’s model catalog places the 2.8 pair in the main Audio section and the 2.6 and 02 pairs under Legacy Models. New integrations should evaluate the 2.8 pair unless a verified compatibility requirement calls for a legacy ID.
| Model | Catalog position | PAYG price verified July 18, 2026 |
|---|---|---|
speech-2.8-turbo | Audio | $60 per 1 million characters |
speech-2.8-hd | Audio | $100 per 1 million characters |
speech-2.6-turbo / speech-02-turbo | Legacy Models | $60 per 1 million characters |
speech-2.6-hd / speech-02-hd | Legacy Models | $100 per 1 million characters |
Prices and account allowances can change independently of this tutorial. Confirm the charge and key entitlement before a production rollout, and use the site’s MiniMax pricing explanation to distinguish Open Platform PAYG billing from other MiniMax plans.
The WebSocket event lifecycle
The connection is a state machine. A successful network handshake is not permission to send synthesis text immediately. Wait for each server event before advancing.
- Open the secure socket. Connect with the Bearer credential and wait for
connected_success. - Configure synthesis. Send
task_startwith a model, voice settings, audio settings, and any supported language or pronunciation controls. - Wait for acceptance. Do not submit text until the server returns
task_started. - Send text. Submit a
task_continueevent. The protocol permits multipletask_continueevents sequentially after the task has started. - Consume audio. Decode each non-empty
data.audiovalue from hexadecimal into bytes. The continued response also carries metadata such as sample rate, size, duration, billed characters, and anis_finalmarker. - Finish the task. After the final result for the queued input, send
task_finish. The server waits for queued work to finish and ends the session. - Confirm completion. Treat
task_finishedas successful completion. Treattask_failedor a non-zerobase_resp.status_codeas an error and close the connection.
Minimal message sequence
SERVER { "event": "connected_success", ... }
CLIENT { "event": "task_start", "model": "speech-2.8-turbo", ... }
SERVER { "event": "task_started", ... }
CLIENT { "event": "task_continue", "text": "Your narration text." }
SERVER { "data": { "audio": "...hex..." }, "is_final": false, ... }
SERVER { "data": { "audio": "...hex..." }, "is_final": true, ... }
CLIENT { "event": "task_finish" }
SERVER { "event": "task_finished", ... }
is_final marks the final continued result; it is not a substitute for the protocol’s finish event. Send task_finish and handle task_finished or socket closure. If a task_failed event arrives, stop writing output, preserve the trace_id for diagnosis, and close the socket.
Node.js streaming client
This server-side example uses the ws package, retains normal TLS certificate verification, writes arriving MP3 bytes to a temporary file, and renames the file only after task_finished. Set the API key in an environment variable; never place it in a frontend bundle.
npm install ws
import WebSocket from "ws";
import { createWriteStream } from "node:fs";
import { rename } from "node:fs/promises";
const API_KEY = process.env.MINIMAX_API_KEY;
const TEXT = process.env.TTS_TEXT ??
"A calm narrator welcomes the listener to the MiniMax WebSocket TTS demo.";
const ENDPOINT = "wss://api.minimax.io/ws/v1/t2a_v2";
const PART_PATH = "minimax-stream.mp3.part";
const FINAL_PATH = "minimax-stream.mp3";
if (!API_KEY) {
throw new Error("Set MINIMAX_API_KEY before running this script.");
}
if ([...TEXT].length > 10_000) {
throw new Error("WebSocket TTS accepts up to 10,000 characters per request.");
}
const output = createWriteStream(PART_PATH, { flags: "w" });
const socket = new WebSocket(ENDPOINT, {
headers: { Authorization: `Bearer ${API_KEY}` },
});
let finishSent = false;
let completed = false;
let failed = false;
let idleTimer;
function resetIdleTimer() {
clearTimeout(idleTimer);
idleTimer = setTimeout(() => {
fail(new Error("No WebSocket event was received before the client timeout."));
}, 125_000);
}
function send(message) {
socket.send(JSON.stringify(message));
}
function checkApiResult(message) {
const statusCode = message.base_resp?.status_code;
if (typeof statusCode === "number" && statusCode !== 0) {
throw new Error(
`MiniMax error ${statusCode}: ${message.base_resp?.status_msg ?? "unknown"}`,
);
}
if (message.event === "task_failed") {
throw new Error(
`TTS task failed. trace_id=${message.trace_id ?? "not returned"}`,
);
}
}
function fail(error) {
if (failed || completed) return;
failed = true;
clearTimeout(idleTimer);
console.error(error.message);
output.destroy();
if (socket.readyState === WebSocket.OPEN) socket.terminate();
process.exitCode = 1;
}
socket.on("open", resetIdleTimer);
socket.on("message", (raw) => {
try {
resetIdleTimer();
const message = JSON.parse(raw.toString("utf8"));
checkApiResult(message);
if (message.event === "connected_success") {
send({
event: "task_start",
model: "speech-2.8-turbo",
language_boost: "English",
voice_setting: {
voice_id: "English_expressive_narrator",
speed: 1,
vol: 1,
pitch: 0,
},
audio_setting: {
sample_rate: 32000,
bitrate: 128000,
format: "mp3",
channel: 1,
},
});
return;
}
if (message.event === "task_started") {
send({ event: "task_continue", text: TEXT });
return;
}
const hexAudio = message.data?.audio;
if (typeof hexAudio === "string" && hexAudio.length > 0) {
output.write(Buffer.from(hexAudio, "hex"));
}
if (message.is_final === true && !finishSent) {
finishSent = true;
send({ event: "task_finish" });
return;
}
if (message.event === "task_finished") {
completed = true;
clearTimeout(idleTimer);
output.end(async () => {
try {
await rename(PART_PATH, FINAL_PATH);
console.log(`Saved ${FINAL_PATH}`);
if (socket.readyState === WebSocket.OPEN) {
socket.close(1000, "TTS complete");
}
} catch (error) {
console.error(`Could not finalize output: ${error.message}`);
process.exitCode = 1;
}
});
}
} catch (error) {
fail(error);
}
});
socket.on("unexpected-response", (_request, response) => {
fail(new Error(`WebSocket handshake failed with HTTP ${response.statusCode}.`));
});
socket.on("error", fail);
socket.on("close", (code) => {
clearTimeout(idleTimer);
if (!completed && !failed) {
fail(new Error(`WebSocket closed before task_finished (code ${code}).`));
}
});
The example streams to disk rather than retaining the full result in memory. To play audio during synthesis, send the same decoded byte chunks to an MP3-capable decoder or media process and keep the file writer as a separate consumer. In a busy service, place a bounded queue between WebSocket messages and downstream consumers so a slow player or storage target cannot cause unbounded buffering.
Python streaming client
The Python version follows the same state transitions and uses the operating system’s trusted certificate store. It does not disable hostname or certificate validation.
python -m pip install websockets
import asyncio
import json
import os
import ssl
from pathlib import Path
import websockets
ENDPOINT = "wss://api.minimax.io/ws/v1/t2a_v2"
PART_PATH = Path("minimax-stream.mp3.part")
FINAL_PATH = Path("minimax-stream.mp3")
def check_api_result(message: dict) -> None:
base = message.get("base_resp") or {}
status_code = base.get("status_code")
if isinstance(status_code, int) and status_code != 0:
raise RuntimeError(
f"MiniMax error {status_code}: {base.get('status_msg', 'unknown')}"
)
if message.get("event") == "task_failed":
trace_id = message.get("trace_id", "not returned")
raise RuntimeError(f"TTS task failed. trace_id={trace_id}")
async def synthesize() -> None:
api_key = os.environ.get("MINIMAX_API_KEY")
text = os.environ.get(
"TTS_TEXT",
"A calm narrator welcomes the listener to the MiniMax WebSocket TTS demo.",
)
if not api_key:
raise RuntimeError("Set MINIMAX_API_KEY before running this script.")
if len(text) > 10_000:
raise ValueError("WebSocket TTS accepts up to 10,000 characters per request.")
headers = {"Authorization": f"Bearer {api_key}"}
ssl_context = ssl.create_default_context()
finish_sent = False
completed = False
try:
async with websockets.connect(
ENDPOINT,
additional_headers=headers,
ssl=ssl_context,
) as websocket:
with PART_PATH.open("wb") as output:
while True:
raw = await asyncio.wait_for(websocket.recv(), timeout=125)
message = json.loads(raw)
check_api_result(message)
event = message.get("event")
if event == "connected_success":
await websocket.send(json.dumps({
"event": "task_start",
"model": "speech-2.8-turbo",
"language_boost": "English",
"voice_setting": {
"voice_id": "English_expressive_narrator",
"speed": 1,
"vol": 1,
"pitch": 0,
},
"audio_setting": {
"sample_rate": 32000,
"bitrate": 128000,
"format": "mp3",
"channel": 1,
},
}))
continue
if event == "task_started":
await websocket.send(json.dumps({
"event": "task_continue",
"text": text,
}))
continue
hex_audio = (message.get("data") or {}).get("audio")
if isinstance(hex_audio, str) and hex_audio:
output.write(bytes.fromhex(hex_audio))
if message.get("is_final") is True and not finish_sent:
finish_sent = True
await websocket.send(json.dumps({"event": "task_finish"}))
continue
if event == "task_finished":
completed = True
break
if not completed:
raise RuntimeError("Connection ended before task_finished.")
PART_PATH.replace(FINAL_PATH)
print(f"Saved {FINAL_PATH}")
except Exception:
PART_PATH.unlink(missing_ok=True)
raise
asyncio.run(synthesize())
Audio chunks, playback, and file integrity
- Decode hexadecimal. Node.js uses
Buffer.from(value, "hex"); Python usesbytes.fromhex(value). - Preserve chunk order. Write or enqueue bytes in the same order in which messages arrive.
- Use a compatible decoder. The examples request MP3, so a player must accept an incremental MP3 byte stream.
- Commit atomically. Write to a
.partfile and rename it only after the success event. This prevents an interrupted result from looking complete. - Do not infer completion from a quiet socket. A timeout, close event, or missing message is not a successful generation result.
- Record metadata. Store
trace_id, returned audio metadata, model ID, voice ID, and input revision alongside the asset, without recording the API key.
Multiple task_continue events
The official protocol permits multiple task_continue events after task_started. Send them sequentially and keep the total request within the documented synchronous limit. A safe application-level strategy is to submit one logical segment, consume its final result, then submit the next segment. Send task_finish only after no further text will be queued.
Do not split text at arbitrary character positions. Preserve sentence and paragraph boundaries, because cutting a sentence can damage prosody and pronunciation. If the material is large enough to need extensive segmentation, the async workflow is usually easier to operate and audit.
Production security checklist
- Keep the key on the server. A browser, mobile bundle, public repository, WordPress page, or application log is not a safe place for an Open Platform credential.
- Verify TLS. Keep certificate and hostname verification enabled. Do not set Node’s
rejectUnauthorizedtofalse, and do not use Python’sCERT_NONE. - Restrict input. Apply length limits before opening a billable task and reject unsupported or control-heavy text.
- Authorize voice use. Use system voices or custom voices for which the operator has the necessary permission and consent.
- Bound concurrency. Respect the account’s T2A limit and place synthesis work behind a queue.
- Separate tenants. Do not allow one customer to select another customer’s private voice ID or download another customer’s audio.
- Protect partial output. Store temporary files outside public directories and delete them after a failed session.
- Minimize logs. Keep status, trace ID, timing, and usage metrics; avoid logging secrets or sensitive narration text.
See the site’s security practices and privacy policy for the independent demo. Those pages describe MiniMax-AI.chat; an API customer remains responsible for its own data flow, retention, access controls, disclosures, and legal basis.
Error handling and recovery
| Symptom | Likely stage | Safe response |
|---|---|---|
No connected_success | Handshake or authentication | Check the server-side key, endpoint, TLS path, and handshake response. Do not expose the key while debugging. |
No task_started | Task configuration | Inspect base_resp, model ID, voice ID, and audio settings before sending text. |
task_failed | Synthesis | Stop consuming output, close the socket, delete the partial file, and retain the trace ID. |
| Unreadable audio | Chunk decoding or format | Confirm hexadecimal decoding, ordered writes, and a decoder that matches the requested format. |
| Connection closes after inactivity | Session lifecycle | The documented idle window is 120 seconds after the last result. Finish promptly and create a fresh session for later work. |
| Rate-limit response | Capacity control | Queue requests and retry with bounded exponential backoff plus jitter. Do not create a rapid reconnect loop. |
Do not automatically replay an unknown task after every disconnect. The public WebSocket reference does not document an idempotency key for this flow, so an indiscriminate retry can create duplicate speech and duplicate charges. Retry only when the application can determine that no completed asset was accepted.
Frequently asked questions
What is the MiniMax WebSocket TTS endpoint?
The documented endpoint is wss://api.minimax.io/ws/v1/t2a_v2. Authenticate the handshake with a Bearer API key from the MiniMax Open Platform.
Can I call it directly from WordPress or browser JavaScript?
Not safely with a private API key. A public browser bundle exposes the credential. Put the MiniMax connection behind a controlled backend that authenticates your user, validates text, enforces quotas, and returns only the required audio stream.
Is data.audio Base64?
No. The WebSocket examples return the audio value as hexadecimal text. Decode it as hex before writing or playing the bytes.
Does is_final close the session?
No. After the queued synthesis result is final and no more text will be sent, submit {"event":"task_finish"} and handle task_finished or the resulting socket closure.
How much text can one WebSocket request contain?
The official synchronous WebSocket guide documents a 10,000-character maximum per request. Route longer jobs to async TTS instead of relying on an undocumented overflow behavior.
Can I use a designed or cloned voice?
The speech API accepts a voice_id. A system, designed, or cloned voice must exist and be available to the same account. Review the separate MiniMax Voice Design API guide before building a custom-voice workflow.
Implementation checklist
- Store
MINIMAX_API_KEYin a server-side secret manager or environment variable. - Connect to the documented
wss://endpoint with TLS verification enabled. - Wait for
connected_successbefore sendingtask_start. - Wait for
task_startedbefore sending anytask_continue. - Decode
data.audioas hex and preserve message order. - Handle
base_resp,task_failed, connection errors, and partial files. - Send
task_finishwhen the queue is complete and confirmtask_finished. - Record usage metadata and trace IDs without logging credentials.
- Load-test with a controlled queue below the verified account allowance.
Official sources
- MiniMax Text to Speech WebSocket API reference
- MiniMax synchronous WebSocket TTS guide
- MiniMax model catalog
- MiniMax Pay as You Go pricing
- MiniMax API rate limits
- MiniMax error code reference
MiniMax-AI.chat is an independent educational website. Product names and trademarks belong to their respective owners.
