MiniMax WebSocket TTS: Streaming Speech With Node.js and Python

MiniMax WebSocket TTS is the synchronous speech route for applications that need audio chunks before an entire utterance has finished synthesizing. A server connects to wss://api.minimax.io/ws/v1/t2a_v2, authenticates with a MiniMax API key, starts a task, submits text, decodes the returned hexadecimal audio, and explicitly finishes the session.

Requires MiniMax Open Platform API. MiniMax-AI.chat is an independent resource and is not operated by MiniMax. The MiniMax-AI.chat demo is text-only and does not expose speech generation or a WebSocket connection.

Last verified: July 18, 2026. Endpoint names, event messages, model IDs, limits, pricing, and rate-limit statements on this page were checked against MiniMax’s official documentation. The examples were syntax-checked, but no paid synthesis request was submitted for this article.

MiniMax WebSocket TTS at a glance

ItemDocumented behaviorImplementation consequence
TransportSecure WebSocketConnect from a trusted backend, not public browser code.
Endpointwss://api.minimax.io/ws/v1/t2a_v2Send the API key in the Authorization: Bearer … handshake header.
Text limitUp to 10,000 characters per synchronous requestUse the asynchronous route for substantially longer material.
Audio deliverydata.audio contains hex-encoded audio chunksDecode with a hex decoder; do not treat the value as Base64.
Client eventstask_start, one or more task_continue events, then task_finishRespect the state machine instead of sending text immediately after opening the socket.
Server eventsconnected_success, task_started, continued results, task_finished, or task_failedGate every client action on the corresponding server acknowledgement.
Idle behaviorThe connection closes if no new event is sent within 120 seconds after the last resultFinish promptly; do not keep an unused session open.
Speech rate limitThe general T2A table lists 60 requests per minuteQueue work and re-check the allowance attached to the production account.

This page deliberately focuses on the WebSocket protocol and streaming client. For a broad API map, see the MiniMax API guide. For long scripts, task polling, downloadable result files, and sentence-level subtitle output, use the separate MiniMax Async TTS guide.

When to use WebSocket TTS

Choose the WebSocket route when an application benefits from receiving audio incrementally: a conversational agent, accessibility reader, interactive language exercise, or server-side narration service. The first playable bytes can be sent to a decoder while later chunks are arriving.

Do not choose it merely because the source text is long. The synchronous WebSocket guide caps a request at 10,000 characters. Async TTS accepts up to 50,000 characters in the direct text field or up to 1,000,000 characters through a supported uploaded file. Those modes solve different problems: WebSocket reduces delivery latency, while async TTS manages long-running jobs and downloadable artifacts.

RequirementBetter routeReason
Play speech while it is generatedWebSocket TTSAudio arrives as incremental hex chunks.
Simple request/response integrationSynchronous HTTP TTSNo persistent bidirectional session is required.
Book, course, podcast script, or batch narrationAsync TTSTask polling, file input, result retrieval, and sentence-level subtitle output are documented.
Direct browser integrationBackend proxy requiredEmbedding a MiniMax API key in client-side JavaScript exposes the credential.

Supported models and billing

The official WebSocket guide lists six model IDs: speech-2.8-hd, speech-2.8-turbo, speech-2.6-hd, speech-2.6-turbo, speech-02-hd, and speech-02-turbo. MiniMax’s model catalog places the 2.8 pair in the main Audio section and the 2.6 and 02 pairs under Legacy Models. New integrations should evaluate the 2.8 pair unless a verified compatibility requirement calls for a legacy ID.

ModelCatalog positionPAYG price verified July 18, 2026
speech-2.8-turboAudio$60 per 1 million characters
speech-2.8-hdAudio$100 per 1 million characters
speech-2.6-turbo / speech-02-turboLegacy Models$60 per 1 million characters
speech-2.6-hd / speech-02-hdLegacy Models$100 per 1 million characters

Prices and account allowances can change independently of this tutorial. Confirm the charge and key entitlement before a production rollout, and use the site’s MiniMax pricing explanation to distinguish Open Platform PAYG billing from other MiniMax plans.

The WebSocket event lifecycle

The connection is a state machine. A successful network handshake is not permission to send synthesis text immediately. Wait for each server event before advancing.

  1. Open the secure socket. Connect with the Bearer credential and wait for connected_success.
  2. Configure synthesis. Send task_start with a model, voice settings, audio settings, and any supported language or pronunciation controls.
  3. Wait for acceptance. Do not submit text until the server returns task_started.
  4. Send text. Submit a task_continue event. The protocol permits multiple task_continue events sequentially after the task has started.
  5. Consume audio. Decode each non-empty data.audio value from hexadecimal into bytes. The continued response also carries metadata such as sample rate, size, duration, billed characters, and an is_final marker.
  6. Finish the task. After the final result for the queued input, send task_finish. The server waits for queued work to finish and ends the session.
  7. Confirm completion. Treat task_finished as successful completion. Treat task_failed or a non-zero base_resp.status_code as an error and close the connection.

Minimal message sequence

SERVER  { "event": "connected_success", ... }
CLIENT  { "event": "task_start", "model": "speech-2.8-turbo", ... }
SERVER  { "event": "task_started", ... }
CLIENT  { "event": "task_continue", "text": "Your narration text." }
SERVER  { "data": { "audio": "...hex..." }, "is_final": false, ... }
SERVER  { "data": { "audio": "...hex..." }, "is_final": true, ... }
CLIENT  { "event": "task_finish" }
SERVER  { "event": "task_finished", ... }

is_final marks the final continued result; it is not a substitute for the protocol’s finish event. Send task_finish and handle task_finished or socket closure. If a task_failed event arrives, stop writing output, preserve the trace_id for diagnosis, and close the socket.

Node.js streaming client

This server-side example uses the ws package, retains normal TLS certificate verification, writes arriving MP3 bytes to a temporary file, and renames the file only after task_finished. Set the API key in an environment variable; never place it in a frontend bundle.

npm install ws
import WebSocket from "ws";
import { createWriteStream } from "node:fs";
import { rename } from "node:fs/promises";

const API_KEY = process.env.MINIMAX_API_KEY;
const TEXT = process.env.TTS_TEXT ??
  "A calm narrator welcomes the listener to the MiniMax WebSocket TTS demo.";
const ENDPOINT = "wss://api.minimax.io/ws/v1/t2a_v2";
const PART_PATH = "minimax-stream.mp3.part";
const FINAL_PATH = "minimax-stream.mp3";

if (!API_KEY) {
  throw new Error("Set MINIMAX_API_KEY before running this script.");
}
if ([...TEXT].length > 10_000) {
  throw new Error("WebSocket TTS accepts up to 10,000 characters per request.");
}

const output = createWriteStream(PART_PATH, { flags: "w" });
const socket = new WebSocket(ENDPOINT, {
  headers: { Authorization: `Bearer ${API_KEY}` },
});

let finishSent = false;
let completed = false;
let failed = false;
let idleTimer;

function resetIdleTimer() {
  clearTimeout(idleTimer);
  idleTimer = setTimeout(() => {
    fail(new Error("No WebSocket event was received before the client timeout."));
  }, 125_000);
}

function send(message) {
  socket.send(JSON.stringify(message));
}

function checkApiResult(message) {
  const statusCode = message.base_resp?.status_code;
  if (typeof statusCode === "number" && statusCode !== 0) {
    throw new Error(
      `MiniMax error ${statusCode}: ${message.base_resp?.status_msg ?? "unknown"}`,
    );
  }
  if (message.event === "task_failed") {
    throw new Error(
      `TTS task failed. trace_id=${message.trace_id ?? "not returned"}`,
    );
  }
}

function fail(error) {
  if (failed || completed) return;
  failed = true;
  clearTimeout(idleTimer);
  console.error(error.message);
  output.destroy();
  if (socket.readyState === WebSocket.OPEN) socket.terminate();
  process.exitCode = 1;
}

socket.on("open", resetIdleTimer);

socket.on("message", (raw) => {
  try {
    resetIdleTimer();
    const message = JSON.parse(raw.toString("utf8"));
    checkApiResult(message);

    if (message.event === "connected_success") {
      send({
        event: "task_start",
        model: "speech-2.8-turbo",
        language_boost: "English",
        voice_setting: {
          voice_id: "English_expressive_narrator",
          speed: 1,
          vol: 1,
          pitch: 0,
        },
        audio_setting: {
          sample_rate: 32000,
          bitrate: 128000,
          format: "mp3",
          channel: 1,
        },
      });
      return;
    }

    if (message.event === "task_started") {
      send({ event: "task_continue", text: TEXT });
      return;
    }

    const hexAudio = message.data?.audio;
    if (typeof hexAudio === "string" && hexAudio.length > 0) {
      output.write(Buffer.from(hexAudio, "hex"));
    }

    if (message.is_final === true && !finishSent) {
      finishSent = true;
      send({ event: "task_finish" });
      return;
    }

    if (message.event === "task_finished") {
      completed = true;
      clearTimeout(idleTimer);
      output.end(async () => {
        try {
          await rename(PART_PATH, FINAL_PATH);
          console.log(`Saved ${FINAL_PATH}`);
          if (socket.readyState === WebSocket.OPEN) {
            socket.close(1000, "TTS complete");
          }
        } catch (error) {
          console.error(`Could not finalize output: ${error.message}`);
          process.exitCode = 1;
        }
      });
    }
  } catch (error) {
    fail(error);
  }
});

socket.on("unexpected-response", (_request, response) => {
  fail(new Error(`WebSocket handshake failed with HTTP ${response.statusCode}.`));
});

socket.on("error", fail);

socket.on("close", (code) => {
  clearTimeout(idleTimer);
  if (!completed && !failed) {
    fail(new Error(`WebSocket closed before task_finished (code ${code}).`));
  }
});

The example streams to disk rather than retaining the full result in memory. To play audio during synthesis, send the same decoded byte chunks to an MP3-capable decoder or media process and keep the file writer as a separate consumer. In a busy service, place a bounded queue between WebSocket messages and downstream consumers so a slow player or storage target cannot cause unbounded buffering.

Python streaming client

The Python version follows the same state transitions and uses the operating system’s trusted certificate store. It does not disable hostname or certificate validation.

python -m pip install websockets
import asyncio
import json
import os
import ssl
from pathlib import Path

import websockets

ENDPOINT = "wss://api.minimax.io/ws/v1/t2a_v2"
PART_PATH = Path("minimax-stream.mp3.part")
FINAL_PATH = Path("minimax-stream.mp3")


def check_api_result(message: dict) -> None:
    base = message.get("base_resp") or {}
    status_code = base.get("status_code")
    if isinstance(status_code, int) and status_code != 0:
        raise RuntimeError(
            f"MiniMax error {status_code}: {base.get('status_msg', 'unknown')}"
        )
    if message.get("event") == "task_failed":
        trace_id = message.get("trace_id", "not returned")
        raise RuntimeError(f"TTS task failed. trace_id={trace_id}")


async def synthesize() -> None:
    api_key = os.environ.get("MINIMAX_API_KEY")
    text = os.environ.get(
        "TTS_TEXT",
        "A calm narrator welcomes the listener to the MiniMax WebSocket TTS demo.",
    )
    if not api_key:
        raise RuntimeError("Set MINIMAX_API_KEY before running this script.")
    if len(text) > 10_000:
        raise ValueError("WebSocket TTS accepts up to 10,000 characters per request.")

    headers = {"Authorization": f"Bearer {api_key}"}
    ssl_context = ssl.create_default_context()
    finish_sent = False
    completed = False

    try:
        async with websockets.connect(
            ENDPOINT,
            additional_headers=headers,
            ssl=ssl_context,
        ) as websocket:
            with PART_PATH.open("wb") as output:
                while True:
                    raw = await asyncio.wait_for(websocket.recv(), timeout=125)
                    message = json.loads(raw)
                    check_api_result(message)

                    event = message.get("event")
                    if event == "connected_success":
                        await websocket.send(json.dumps({
                            "event": "task_start",
                            "model": "speech-2.8-turbo",
                            "language_boost": "English",
                            "voice_setting": {
                                "voice_id": "English_expressive_narrator",
                                "speed": 1,
                                "vol": 1,
                                "pitch": 0,
                            },
                            "audio_setting": {
                                "sample_rate": 32000,
                                "bitrate": 128000,
                                "format": "mp3",
                                "channel": 1,
                            },
                        }))
                        continue

                    if event == "task_started":
                        await websocket.send(json.dumps({
                            "event": "task_continue",
                            "text": text,
                        }))
                        continue

                    hex_audio = (message.get("data") or {}).get("audio")
                    if isinstance(hex_audio, str) and hex_audio:
                        output.write(bytes.fromhex(hex_audio))

                    if message.get("is_final") is True and not finish_sent:
                        finish_sent = True
                        await websocket.send(json.dumps({"event": "task_finish"}))
                        continue

                    if event == "task_finished":
                        completed = True
                        break

        if not completed:
            raise RuntimeError("Connection ended before task_finished.")
        PART_PATH.replace(FINAL_PATH)
        print(f"Saved {FINAL_PATH}")
    except Exception:
        PART_PATH.unlink(missing_ok=True)
        raise


asyncio.run(synthesize())

Audio chunks, playback, and file integrity

  • Decode hexadecimal. Node.js uses Buffer.from(value, "hex"); Python uses bytes.fromhex(value).
  • Preserve chunk order. Write or enqueue bytes in the same order in which messages arrive.
  • Use a compatible decoder. The examples request MP3, so a player must accept an incremental MP3 byte stream.
  • Commit atomically. Write to a .part file and rename it only after the success event. This prevents an interrupted result from looking complete.
  • Do not infer completion from a quiet socket. A timeout, close event, or missing message is not a successful generation result.
  • Record metadata. Store trace_id, returned audio metadata, model ID, voice ID, and input revision alongside the asset, without recording the API key.

Multiple task_continue events

The official protocol permits multiple task_continue events after task_started. Send them sequentially and keep the total request within the documented synchronous limit. A safe application-level strategy is to submit one logical segment, consume its final result, then submit the next segment. Send task_finish only after no further text will be queued.

Do not split text at arbitrary character positions. Preserve sentence and paragraph boundaries, because cutting a sentence can damage prosody and pronunciation. If the material is large enough to need extensive segmentation, the async workflow is usually easier to operate and audit.

Production security checklist

  • Keep the key on the server. A browser, mobile bundle, public repository, WordPress page, or application log is not a safe place for an Open Platform credential.
  • Verify TLS. Keep certificate and hostname verification enabled. Do not set Node’s rejectUnauthorized to false, and do not use Python’s CERT_NONE.
  • Restrict input. Apply length limits before opening a billable task and reject unsupported or control-heavy text.
  • Authorize voice use. Use system voices or custom voices for which the operator has the necessary permission and consent.
  • Bound concurrency. Respect the account’s T2A limit and place synthesis work behind a queue.
  • Separate tenants. Do not allow one customer to select another customer’s private voice ID or download another customer’s audio.
  • Protect partial output. Store temporary files outside public directories and delete them after a failed session.
  • Minimize logs. Keep status, trace ID, timing, and usage metrics; avoid logging secrets or sensitive narration text.

See the site’s security practices and privacy policy for the independent demo. Those pages describe MiniMax-AI.chat; an API customer remains responsible for its own data flow, retention, access controls, disclosures, and legal basis.

Error handling and recovery

SymptomLikely stageSafe response
No connected_successHandshake or authenticationCheck the server-side key, endpoint, TLS path, and handshake response. Do not expose the key while debugging.
No task_startedTask configurationInspect base_resp, model ID, voice ID, and audio settings before sending text.
task_failedSynthesisStop consuming output, close the socket, delete the partial file, and retain the trace ID.
Unreadable audioChunk decoding or formatConfirm hexadecimal decoding, ordered writes, and a decoder that matches the requested format.
Connection closes after inactivitySession lifecycleThe documented idle window is 120 seconds after the last result. Finish promptly and create a fresh session for later work.
Rate-limit responseCapacity controlQueue requests and retry with bounded exponential backoff plus jitter. Do not create a rapid reconnect loop.

Do not automatically replay an unknown task after every disconnect. The public WebSocket reference does not document an idempotency key for this flow, so an indiscriminate retry can create duplicate speech and duplicate charges. Retry only when the application can determine that no completed asset was accepted.

Frequently asked questions

What is the MiniMax WebSocket TTS endpoint?

The documented endpoint is wss://api.minimax.io/ws/v1/t2a_v2. Authenticate the handshake with a Bearer API key from the MiniMax Open Platform.

Can I call it directly from WordPress or browser JavaScript?

Not safely with a private API key. A public browser bundle exposes the credential. Put the MiniMax connection behind a controlled backend that authenticates your user, validates text, enforces quotas, and returns only the required audio stream.

Is data.audio Base64?

No. The WebSocket examples return the audio value as hexadecimal text. Decode it as hex before writing or playing the bytes.

Does is_final close the session?

No. After the queued synthesis result is final and no more text will be sent, submit {"event":"task_finish"} and handle task_finished or the resulting socket closure.

How much text can one WebSocket request contain?

The official synchronous WebSocket guide documents a 10,000-character maximum per request. Route longer jobs to async TTS instead of relying on an undocumented overflow behavior.

Can I use a designed or cloned voice?

The speech API accepts a voice_id. A system, designed, or cloned voice must exist and be available to the same account. Review the separate MiniMax Voice Design API guide before building a custom-voice workflow.

Implementation checklist

  • Store MINIMAX_API_KEY in a server-side secret manager or environment variable.
  • Connect to the documented wss:// endpoint with TLS verification enabled.
  • Wait for connected_success before sending task_start.
  • Wait for task_started before sending any task_continue.
  • Decode data.audio as hex and preserve message order.
  • Handle base_resp, task_failed, connection errors, and partial files.
  • Send task_finish when the queue is complete and confirm task_finished.
  • Record usage metadata and trace IDs without logging credentials.
  • Load-test with a controlled queue below the verified account allowance.

Official sources

MiniMax-AI.chat is an independent educational website. Product names and trademarks belong to their respective owners.