MiniMax Speech 2.8 Multilingual Benchmark: Results from 48 HD and Turbo Calls

We tested Speech 2.8 HD and Turbo across eight languages with three repetitions per model-language cell. This objective API benchmark reports delivery reliability, latency, real-time factor, WAV signal integrity, and cost—not human-rated naturalness.

We sent 48 scored text-to-speech requests through the MiniMax Speech 2.8 HTTP API on July 28, 2026: eight languages, two models, and three repetitions per language-model pair. Every scored call returned a valid 32 kHz mono WAV file and passed our predefined mechanical signal checks. Speech 2.8 Turbo recorded the lower overall median completion latency and a 40% lower list-price estimate than Speech 2.8 HD for the same API-reported character usage.

Important scope note: this is an objective API delivery and WAV-integrity benchmark. We did not use speech recognition, native-speaker ratings, MOS panels, or preference voting, so these results do not establish which model sounds more natural or pronounces each language more accurately.

MiniMax Speech 2.8 test results at a glance

  • 48/48 scored calls delivered successfully. Two additional English warm-up calls also succeeded but were excluded from every scored result.
  • 48/48 files passed format and signal gates. Every scored output decoded as 32 kHz, mono, 16-bit WAV audio.
  • Turbo had the lower overall median completion latency: 3,566 ms versus 3,694 ms for HD, a 128 ms or 3.5% difference in this run.
  • Turbo also had the lower overall p95: 3,961.4 ms versus 4,329.2 ms for HD, an 8.5% difference.
  • Median real-time factor was 0.193 for Turbo and 0.206 for HD. Both models generated the complete non-streaming response in roughly one fifth of the resulting audio duration.
  • Scored list-price estimate: $0.28512 for Turbo and $0.47520 for HD, based on 4,752 API-reported characters per model. The 48 scored calls totaled $0.76032.

Those timing differences describe one sequential run from one account, machine, network, and API route. Turbo produced the lower median latency in six of the eight language cells; HD was lower for Mandarin Chinese and French. With only three repetitions per cell, the language-level differences should be treated as observations, not universal performance guarantees.

Overall HD vs Turbo results

MetricSpeech 2.8 HDSpeech 2.8 Turbo
Scored calls2424
Successful delivery24/24 (100%)24/24 (100%)
Format pass24/24 (100%)24/24 (100%)
Signal pass24/24 (100%)24/24 (100%)
Median completion latency3,694 ms3,566 ms
p95 completion latency4,329.2 ms3,961.4 ms
Median real-time factor0.2059540.193212
p95 real-time factor0.2346860.242122
Mean output duration17.821 s18.243 s
API-reported usage4,752 characters4,752 characters
List-price estimate$0.47520$0.28512
Results from protocol M28-M8B-48-v1.0. Latency is time to the final byte of a non-streaming response, not streaming time to first audio.
Median non-streaming completion latency by language for MiniMax Speech 2.8 HD and Turbo, based on three scored calls per cell.
Median completion latency by language. Lower is better, but three repetitions per cell are not enough to characterize long-term service variance.
Median real-time factor by language for MiniMax Speech 2.8 HD and Turbo, based on three scored calls per cell.
Median real-time factor by language. An RTF below 1 means the full output arrived faster than the duration of the generated audio.

Results by language

Each row below represents six scored calls: three HD and three Turbo. The two models received the same localized text and the same system voice within a language. Different languages used different system voices and different text lengths, so this table is suitable for comparing HD with Turbo within a language—not for ranking languages or voices against one another.

LanguageHD p50 latencyTurbo p50 latencyHD p50 RTFTurbo p50 RTFDelivery + integrityHD / Turbo cost
English3,492 ms3,215 ms0.2281320.1941216/6 passed$0.06420 / $0.03852
Mandarin Chinese3,397 ms3,769 ms0.1947230.1932796/6 passed$0.04200 / $0.02520
Japanese3,922 ms3,740 ms0.2043890.1858846/6 passed$0.04470 / $0.02682
Spanish4,285 ms3,733 ms0.2045930.1682006/6 passed$0.06510 / $0.03906
French3,478 ms3,784 ms0.2015140.2129386/6 passed$0.06810 / $0.04086
German3,442 ms3,408 ms0.2192810.2164756/6 passed$0.07410 / $0.04446
Arabic3,707 ms3,605 ms0.1865040.1731786/6 passed$0.05850 / $0.03510
Hindi3,413 ms3,217 ms0.2119410.2009146/6 passed$0.05850 / $0.03510
Cell medians use three repetitions. Cost uses API-reported usage characters and the official list prices checked on July 28, 2026.

What “100% WAV integrity” means here

Our integrity result is deliberately mechanical. A call had to return HTTP 200, an API status code of 0, decodable audio, and a WAV file matching the frozen format. The decoded file then had to pass limits for duration, metadata agreement, clipping, DC offset, and silence. Passing those checks means the output was delivered in the requested structure without an obvious machine-detectable signal failure. It does not mean the speech was linguistically perfect or subjectively natural.

  • Requested and observed format: WAV, 32,000 Hz, mono, 16-bit.
  • Minimum allowed duration: 0.5 seconds.
  • API metadata versus decoded duration tolerance: the greater of 100 ms or 1%.
  • Maximum clipping fraction: 0.001, measured at an absolute normalized amplitude threshold of 0.999.
  • Maximum absolute DC offset: 0.01.
  • Maximum leading, trailing, or internal silence: 2 seconds, using 100 ms windows and a -50 dBFS silence threshold.

Across the 48 scored files, the largest metadata-to-decoded-duration difference was about 1.03 ms. The longest observed leading silence was 100 ms, trailing silence was 269 ms, and internal silence was 900 ms. HD’s largest clipping fraction was 0.000270508—below the frozen 0.001 gate—while Turbo’s maximum was zero at the same threshold. These values are engineering diagnostics, not listening scores.

English audio samples from the run

For transparency, these players use repetition 1 from the scored English cell. They are examples, not a controlled preference test. Use headphones if you want to inspect the outputs yourself, and do not generalize from one voice or one sentence.

Speech 2.8 HD — English, repetition 1
Speech 2.8 Turbo — English, repetition 1

How we ran the 48-call benchmark

We froze protocol M28-M8B-48-v1.0 before execution. The run used MiniMax’s documented HTTP text-to-audio endpoint, POST https://api.minimax.io/v1/t2a_v2, from a Windows x64 machine with Node.js 24.14.0. It started at 19:26:54 UTC and completed at 19:30:00 UTC on July 28, 2026.

Call matrix

  1. Eight languages: English, Mandarin Chinese, Japanese, Spanish, French, German, Arabic, and Hindi.
  2. Two models: speech-2.8-hd and speech-2.8-turbo.
  3. One frozen localized prompt per language.
  4. Three scored repetitions for every language-model pair: 8 × 2 × 3 = 48 scored calls.
  5. One English warm-up per model before scoring: two calls, explicitly excluded.

Scored order was randomized once with seed 20260728. Requests were sequential, with a minimum 1.1-second interval between starts. We allowed no automatic retries, and every attempt remained in the denominator. The timeout was 120 seconds per call.

Frozen request settings

SettingValue
Streamingfalse
Output transportHex audio decoded locally
AudioWAV, 32 kHz, mono
Voice speed / volume / pitch1 / 1 / 0
Language controlExplicit language_boost for every request
Subtitlesfalse
Omitted controlsEmotion, pronunciation dictionary, pause or sound tags, voice modification, and timbre weights

The system voices were English_expressive_narrator, Chinese (Mandarin)_News_Anchor, Japanese_IntellectualSenior, Spanish_Narrator, French_MaleNarrator, German_FriendlyMan, Arabic_CalmWoman, and hindi_male_1_v2. MiniMax documents model availability and capabilities in its model guide and publishes system voice IDs.

Frozen text design

Each localized prompt used the same practical structure: a museum opening time, quantities and colored objects, a quoted question about a train platform, and a short sequence to repeat. That combination exercises prose, numbers, punctuation, and a direct question. The English source was:

On Thursday, the museum opens at nine thirty. Bring twelve red cards, two clean keys, and the yellow notebook. Before you leave, ask, “Did the train arrive at platform four?” Then repeat: north, cedar, twenty-four.

The evidence bundle contains all eight exact UTF-8 prompts, voice IDs, language boosts, request payloads, and the frozen randomized order. Because the text lengths and voices differ across languages, we did not use the data to claim that one language is “faster” or “better” than another.

How latency, real-time factor, and cost were calculated

Completion latency

Latency starts immediately before the HTTP request and stops after the final response byte. Since stream=false, this is full-response latency, sometimes called time to last byte. It is not time to first audio, and the results should not be used as a streaming-latency claim. For streaming implementation details, see our MiniMax WebSocket TTS guide.

Real-time factor

Real-time factor equals completion latency divided by decoded audio duration. For example, an RTF of 0.20 means the non-streaming response completed in about 20% of the generated clip’s playing time. RTF helps normalize timing when the output clips are not identical in duration.

List-price cost

We used each response’s extra_info.usage_characters, not a local character estimate, and applied the official pay-as-you-go prices checked on July 28, 2026: $100 per million characters for Speech 2.8 HD and $60 per million for Speech 2.8 Turbo.

The frozen prompt files contain 4,452 Unicode code points for one complete model pass, while the API returned 4,752 usage characters per model across the scored calls. We did not speculate about the reason for that 300-character difference; we used the returned usage value for the final $0.76032 estimate. The two excluded warm-ups added an estimated $0.03424. These are list-price calculations, not an account invoice, and pricing or billing rules can change.

Download the reproducibility evidence

The downloadable bundle contains the frozen protocol, eight prompts, randomized run order, runner and analysis source, sanitized request and response records, all 50 WAV files, per-call metrics, CSV summaries, charts, provenance, and SHA-256 checksums. The 50 files comprise 48 scored outputs plus two excluded warm-ups.

Package maintenance — verified July 30, 2026: This ZIP was rebuilt with portable POSIX paths and a verified internal checksum manifest. The published Speech 2.8 result is unchanged. ZIP SHA-256: 622ee91efc5a45f4c01d155649255de796725e413c99e6d0433cc7cfd7b43299.

Security was part of the artifact design. Authorization headers were never logged; response audio hex was decoded and replaced in JSON by its filename, byte length, and SHA-256; and the API key was stored outside the protocol directory. Before packaging, an automated scan checked 156 text artifacts and found zero Bearer tokens, zero unredacted audio-hex fields, and zero likely secret fields.

For broader implementation context, compare our MiniMax Speech 2.8 model guide, MiniMax API documentation hub, pricing guide, multilingual voice localization guide, and MiniMax Audio guide.

What these results mean for choosing HD or Turbo

Choose Turbo for an initial production trial when cost and full-response speed matter most. In this specific run, it matched HD’s 100% delivery and mechanical integrity rates, had slightly lower aggregate latency and RTF medians, and cost 40% less at the published per-character rates.

Do not choose a model for voice quality from this benchmark alone. HD may be intended for a different quality-versus-speed tradeoff, but our measurements cannot verify naturalness, pronunciation, accent, emotion, or listener preference. A responsible selection process should run both models on your own scripts and voices, then add blinded native-speaker review alongside the delivery and latency checks used here.

Limitations

  • We tested eight languages, not MiniMax’s complete supported-language set.
  • Each language used one text and one system voice, with only three scored repetitions per model.
  • There was no ASR transcription, WER/CER scoring, native-speaker panel, MOS test, pronunciation rubric, speaker-similarity test, or blinded preference test.
  • Calls were sequential. We did not test concurrency, sustained load, rate-limit behavior, streaming time to first audio, or WebSocket performance.
  • Timing reflects one Windows machine, account, network path, API region, and three-minute run window on July 28, 2026.
  • Cross-language speed and signal-level comparisons are confounded by different text lengths, scripts, voices, and output durations.
  • Mechanical signal gates detect malformed or extreme output, not subtle audible defects or linguistic errors.
  • List-price estimates use the documentation and API usage fields available on the test date; actual invoices and future pricing may differ.

Frequently asked questions

Did every MiniMax Speech 2.8 call succeed?

Yes. All 48 scored calls returned HTTP 200, API status 0, valid requested-format audio, and passed the frozen signal gates. Both excluded English warm-ups also succeeded. This is a 50-call observation, not a long-term availability guarantee.

Was Speech 2.8 Turbo faster than HD?

At the aggregate level in this run, yes: Turbo’s median full-response latency was 3,566 ms versus 3,694 ms for HD, and its p95 was 3,961.4 ms versus 4,329.2 ms. Turbo had the lower cell median in six of eight languages, while HD was lower for Mandarin Chinese and French.

Was Speech 2.8 Turbo cheaper?

Yes at the official list prices checked on July 28, 2026. Both models reported 4,752 scored usage characters. That produced an estimate of $0.28512 for Turbo and $0.47520 for HD—a 40% reduction for Turbo. This is not a guarantee of your final invoice.

Does 100% signal integrity mean perfect pronunciation?

No. It means the files decoded correctly and stayed within predefined engineering limits for format, duration agreement, clipping, DC offset, and silence. Only a separate linguistic evaluation with suitable human or validated ASR scoring could address pronunciation accuracy.

Did this benchmark measure naturalness?

No. We intentionally make no claim about naturalness, accent, emotional expression, voice similarity, or listener preference. The embedded English samples are provided for transparency, not as a scored listening test.

Is the latency result time to first audio?

No. The requests were non-streaming, so latency is time to the final response byte. A streaming or WebSocket benchmark needs a different protocol that records time to first playable audio.

Can I reproduce the test?

Yes. Download the evidence bundle for the exact prompts, manifest, randomized order, source, sanitized call artifacts, CSV data, WAV files, and checksums. Re-running will incur API charges and may produce different timing or audio because service conditions and model behavior can change.

More on speech: see the MiniMax Audio hub for voices, voice cloning, text-to-speech pricing and which speech model to choose.