MiniMax Music Cover API: One-Step and Two-Step Workflows

The MiniMax Music Cover API creates a different rendition of reference audio while using a prompt to describe the target style. It supports a quick one-step request, where MiniMax extracts the source lyrics automatically, and a two-step workflow, where an application first extracts reusable audio features and structured lyrics, lets a user edit those lyrics, and then generates the cover.

Independent-site notice: MiniMax-AI.chat is an independent educational website. It is not owned by, operated by, endorsed by, or affiliated with MiniMax. The API, models, account access, billing, moderation, and generated audio are provided by MiniMax through its official Open Platform.

Last verified: August 20, 2026. Endpoint names, model IDs, request fields, input limits, response fields, and rate-limit statements were checked against the official Music Generation API, Music Cover Preprocess API, and music generation guide. MiniMax says paid Music Generation access is closed to new users; existing paying users may continue. Because Music Cover runs through Music Generation, new users should not plan a hosted Cover integration. The free ID music-cover-free is discontinued. MiniMax Audio and the downloadable Music 3 checkpoint are the official alternatives named in the notice, but neither is documented as a drop-in replacement for this reference-audio Cover workflow.

Evidence status: The examples below were reviewed against the documented JSON schemas and checked for language syntax. No signed-in MiniMax account, copyrighted recording, paid cover generation, or audio-quality test was used for this article. The page therefore explains the documented integration without claiming a listening result, generation time, or style-transfer score.

Rights warning: A working endpoint does not grant permission to reproduce a composition, lyrics, recording, performer identity, or voice. Use reference audio and lyrics only when you own the necessary rights or have permission for the intended territory and distribution channel. Review MiniMax’s applicable terms and obtain specialist advice for commercial releases.

MiniMax cover API endpoints at a glance

PurposeEndpointModel valueWhat it returns
One-step coverPOST https://api.minimax.io/v1/music_generationmusic-cover for an eligible existing paying account; music-cover-free is discontinuedGenerated audio and status metadata
Preprocess source audioPOST https://api.minimax.io/v1/music_cover_preprocessmusic-cover onlyFeature ID, ASR lyrics, structure analysis, duration, and trace ID
Two-step cover generationPOST https://api.minimax.io/v1/music_generationmusic-cover for an eligible existing paying account; music-cover-free is discontinuedGenerated audio using the feature ID and edited lyrics

Both routes use an API key in Authorization: Bearer YOUR_KEY and JSON with Content-Type: application/json. Keep the key on a server. A browser bundle, public repository, WordPress front-end script, or mobile binary is not a safe place for a permanent MiniMax credential. The MiniMax API key and base URL guide covers credential setup and regional host differences.

One-step or two-step: which workflow should you use?

DecisionOne-step coverTwo-step cover
Reference inputSend audio_url or audio_base64 directly to Music Generation.Send the audio to Cover Preprocess, then send its cover_feature_id to Music Generation.
LyricsOmit lyrics to let ASR extract them, or provide 10–1,000 characters.Review and edit formatted_lyrics; lyrics are required in the generation request.
Best fitFast prototypes and unchanged source wording.Corrections, localization, redaction, approved lyric replacements, or a review interface.
Extra dataNo separate structure-analysis response.Receives segment labels and timestamps as a JSON string.
ExpiryNo feature ID is created.The feature ID is valid for 24 hours.

Two-step mode is not a general-purpose stem separator or source-file storage service. The preprocessing response exposes a feature identifier and analysis metadata for the cover workflow; it does not document vocals, drums, bass, or instrumental stems as downloadable assets. It also does not create a File Management API record.

Reference-audio requirements

  • Duration: 6 seconds through 6 minutes.
  • Maximum size: 50 MB.
  • Format: common audio formats such as MP3, WAV, and FLAC are documented.
  • Input choice: provide exactly one of audio_url or audio_base64.
  • Two-step generation: provide cover_feature_id instead of either raw-audio field.

For a URL input, use an HTTPS URL that MiniMax can fetch without browser cookies, a VPN, or an interactive login. If access is temporary, give the URL enough lifetime for the request to begin and avoid exposing a broadly reusable link. For Base64, send the encoded audio string in JSON; do not include it in application logs. Base64 increases the request body’s size, so URL delivery is often easier for a large track.

Validate duration and file size before calling the API. Early validation gives a user a clear message, avoids transferring an unusable payload, and prevents repeated retries for a request that cannot satisfy the schema.

One-step cover request fields

FieldRequirement for a coverImplementation note
modelUse music-cover only if an existing paying account still has access; music-cover-free is discontinued.Do not substitute music-3.0; it is a text-to-music ID, not a Cover model.
promptRequired, 10–300 characters.Describe the target style, mood, instrumentation, vocal treatment, and production character.
audio_url / audio_base64Exactly one is required when no feature ID is used.Never send both.
lyricsOptional, 10–1,000 characters.If omitted, MiniMax extracts lyrics from the reference audio with ASR.
streamOptional; default false.Streaming supports only hexadecimal output.
output_formathex or url; default hex.Documented URL output expires after 24 hours.
audio_settingOptional output configuration.Sample rates: 16000, 24000, 32000, 44100; bitrates: 32000, 64000, 128000, 256000; formats: MP3, WAV, PCM.

Node.js: one-step cover with automatic lyric extraction

This server-side example requests non-streaming hexadecimal MP3 output. Hex is used because its response contract is explicit: the audio is returned in data.audio. Replace the placeholder reference URL with audio you are authorized to process.

import { writeFile } from "node:fs/promises";

const API_KEY = process.env.MINIMAX_API_KEY;
if (!API_KEY) throw new Error("Set MINIMAX_API_KEY on the server.");

async function postMiniMax(path, body) {
  const response = await fetch("https://api.minimax.io" + path, {
    method: "POST",
    headers: {
      Authorization: "Bearer " + API_KEY,
      "Content-Type": "application/json"
    },
    body: JSON.stringify(body)
  });

  const raw = await response.text();
  let result;
  try {
    result = raw ? JSON.parse(raw) : {};
  } catch {
    throw new Error("MiniMax returned non-JSON data (HTTP " + response.status + ").");
  }

  if (!response.ok) {
    throw new Error("MiniMax HTTP " + response.status + ": " + raw);
  }

  const code = result?.base_resp?.status_code;
  if (typeof code === "number" && code !== 0) {
    throw new Error(
      "MiniMax error " + code + ": " +
      (result?.base_resp?.status_msg || "No message")
    );
  }

  return result;
}

const result = await postMiniMax("/v1/music_generation", {
  model: "music-cover",
  audio_url: "https://media.example.com/authorized-reference.mp3",
  prompt:
    "Intimate acoustic folk, fingerpicked guitar, soft percussion, " +
    "restrained vocal, warm room ambience",
  stream: false,
  output_format: "hex",
  audio_setting: {
    sample_rate: 44100,
    bitrate: 256000,
    format: "mp3"
  }
});

if (result?.data?.status !== 2) {
  throw new Error("Music generation did not report completed status 2.");
}

const audioHex = result?.data?.audio;
if (typeof audioHex !== "string" || !/^[0-9a-f]+$/i.test(audioHex)) {
  throw new Error("The response did not contain hexadecimal audio.");
}

await writeFile("authorized-acoustic-cover.mp3", Buffer.from(audioHex, "hex"));

console.log({
  traceId: result.trace_id,
  durationMs: result?.extra_info?.music_duration,
  sampleRate: result?.extra_info?.music_sample_rate,
  bytes: result?.extra_info?.music_size
});

Omitting lyrics instructs the cover workflow to use ASR extraction. That is convenient, but ASR can mishear names, multilingual lines, harmonies, or dense mixes. If wording matters, use the two-step flow and require a human review before generation.

What Cover Preprocess returns

Response fieldMeaningWhat your application should do
cover_feature_idIdentifier for extracted audio features; valid for 24 hours.Store it with an expiry time and use it only in a cover generation request.
formatted_lyricsASR lyrics with structural labels such as Verse and Chorus.Show an editable copy and preserve the original extraction for audit or rollback.
structure_resultA JSON string containing segment labels and start/end times.Parse it once; do not treat the outer string as an already parsed object.
audio_durationReference duration in seconds.Use it for display and validation, not as a promise of output duration.
trace_idRequest-tracking identifier.Retain it for support without recording the API key or full source audio.

The official schema says identical audio content returns the same feature ID through MD5-based deduplication. That behavior can reduce duplicate preprocessing, but it is not a permanent cache: the identifier still expires after 24 hours. Do not use the feature ID as proof of ownership, a public asset ID, or a durable database key.

Node.js: preprocess, inspect, edit, and generate

The following example reuses the postMiniMax helper from the one-step example. In a real application, pause after preprocessing, present formatted_lyrics to an authorized reviewer, and submit only the approved text.

const sourceUrl = "https://media.example.com/authorized-reference.mp3";

const prepared = await postMiniMax("/v1/music_cover_preprocess", {
  model: "music-cover", // The preprocess endpoint accepts this exact ID.
  audio_url: sourceUrl
});

if (!prepared.cover_feature_id || !prepared.formatted_lyrics) {
  throw new Error("The preprocess response is missing features or lyrics.");
}

let structure;
try {
  structure = JSON.parse(prepared.structure_result);
} catch {
  throw new Error("structure_result was not valid JSON text.");
}

console.log({
  featureId: prepared.cover_feature_id,
  sourceDurationSeconds: prepared.audio_duration,
  segments: structure?.segments,
  extractedLyrics: prepared.formatted_lyrics
});

// Replace this value with lyrics approved in your review interface.
const approvedLyrics = `[Verse]
Morning finds the open road
Quiet wheels and lighter loads

[Chorus]
We carry home in every mile
We cross the rain and keep the light`;

if (approvedLyrics.length < 10 || approvedLyrics.length > 1000) {
  throw new Error("Cover lyrics must contain 10–1,000 characters.");
}

const cover = await postMiniMax("/v1/music_generation", {
  model: "music-cover",
  cover_feature_id: prepared.cover_feature_id,
  lyrics: approvedLyrics,
  prompt:
    "Cinematic Americana, brushed drums, resonant guitar, warm ensemble " +
    "vocal, gradual lift into the chorus",
  stream: false,
  output_format: "hex",
  audio_setting: {
    sample_rate: 44100,
    bitrate: 256000,
    format: "mp3"
  }
});

if (cover?.data?.status !== 2 || typeof cover?.data?.audio !== "string") {
  throw new Error("The cover response did not contain completed audio.");
}

await writeFile("approved-two-step-cover.mp3", Buffer.from(cover.data.audio, "hex"));

Notice what the second request does not contain: there is no audio_url and no audio_base64. Those fields are mutually exclusive with cover_feature_id. When a feature ID is used, lyrics becomes required and must contain 10–1,000 characters.

Python: compact two-step cover example

import os
from pathlib import Path
import requests

API_KEY = os.environ.get("MINIMAX_API_KEY")
if not API_KEY:
    raise RuntimeError("Set MINIMAX_API_KEY on the server.")

HEADERS = {
    "Authorization": f"Bearer {API_KEY}",
    "Content-Type": "application/json",
}

def post_minimax(path, payload):
    response = requests.post(
        "https://api.minimax.io" + path,
        headers=HEADERS,
        json=payload,
        timeout=180,
    )
    response.raise_for_status()
    result = response.json()
    code = result.get("base_resp", {}).get("status_code")
    if code not in (None, 0):
        message = result.get("base_resp", {}).get("status_msg", "No message")
        raise RuntimeError(f"MiniMax error {code}: {message}")
    return result

prepared = post_minimax(
    "/v1/music_cover_preprocess",
    {
        "model": "music-cover",
        "audio_url": "https://media.example.com/authorized-reference.mp3",
    },
)

# In production, replace this with human-reviewed formatted_lyrics.
approved_lyrics = prepared["formatted_lyrics"]

cover = post_minimax(
    "/v1/music_generation",
    {
        "model": "music-cover",
        "cover_feature_id": prepared["cover_feature_id"],
        "lyrics": approved_lyrics,
        "prompt": "Dream pop, airy guitars, soft drums, intimate vocal, wide mix",
        "output_format": "hex",
        "audio_setting": {
            "sample_rate": 44100,
            "bitrate": 256000,
            "format": "mp3",
        },
    },
)

if cover.get("data", {}).get("status") != 2:
    raise RuntimeError("Music generation did not complete.")

Path("two-step-cover.mp3").write_bytes(bytes.fromhex(cover["data"]["audio"]))

How to write a cover style prompt

The prompt describes the target cover style, not the source URL or rights status. A useful prompt is concrete enough to guide arrangement and production but short enough to remain within the documented 10–300-character cover range.

  • Genre or arrangement: acoustic folk, jazz trio, dream pop, orchestral ballad.
  • Instrumentation: brushed drums, upright bass, nylon guitar, analog pads.
  • Mood and energy: restrained, triumphant, nocturnal, playful, gradual build.
  • Vocal direction: intimate ensemble, clear lead, soft harmonies. Do not request an identifiable person’s voice without authorization.
  • Production character: dry room, wide mix, vintage tape color, polished broadcast mix.

Example: “Neo-soul trio, warm electric piano, rounded bass, brushed drums, intimate lead vocal, restrained verses, wider final chorus.” Keep the prompt focused on audible properties. Repeating the same adjective or adding unrelated story details does not provide a reliable control mechanism.

Output, storage, and observability

  • Hex output: decode data.audio using a hexadecimal decoder, not Base64.
  • URL output: download it promptly; MiniMax documents a 24-hour expiry.
  • Status: the response schema defines 1 as in progress and 2 as completed.
  • Metadata: extra_info can include duration in milliseconds, sample rate, channels, bitrate, and byte size.
  • Support: store trace_id with your internal job ID.
  • Security: redact API keys, full Base64 audio, temporary signed URLs, and copyrighted lyrics from logs.

The API supports stream: true, but streaming is limited to hex output. Non-streaming is easier for a first integration because one response can be validated and decoded. If you adopt streaming, build a byte-safe accumulator, confirm event framing from an authorized test request, and do not assume that each network chunk is a complete audio frame.

Rate limits and cost planning

The schema still lists music-cover with 120 RPM and the rate-limit table lists 20 concurrent connections, but MiniMax’s August 20 notice closes new paid Music Generation access. Treat those limits as conditional values for an eligible existing paying account. music-cover-free is discontinued; its historical 3-RPM schema value is not current availability. Apply queueing rather than launching an unbounded Promise.all.

The official guide describes Cover Preprocess as free of charge. However, the public pay-as-you-go music table did not show a separate per-cover generation price on the verification date. Do not infer the Music 3.0 song price for music-cover. An eligible existing paying user should check the MiniMax console and applicable account terms before presenting a cost to users. A successful free preprocess response alone does not prove that the account can complete cover generation. The independent MiniMax pricing guide can help you distinguish API balance, Credits, and Token Plan access.

Common integration failures

SymptomLikely causeCorrection
2013 or invalid parametersWrong model, out-of-range prompt or lyrics, two audio inputs, or audio plus feature ID.Validate the workflow’s exact field matrix before sending.
1002Rate limit triggered.Queue requests and retry with capped exponential backoff and jitter.
1004 or 2049Authentication or API-key problem.Confirm the Bearer header, environment, key type, and active account.
1008Insufficient balance.Inspect billing and whether the chosen key has access to the resource.
1026Content was flagged.Do not retry unchanged content in a loop; send it to your review policy.
Feature request fails hours laterThe 24-hour feature ID expired.Preprocess the authorized source again, then regenerate.
Lyrics sound wrongOne-step ASR misheard the source.Use two-step mode and correct the extracted text before generation.

MiniMax can return an HTTP error and can also return HTTP 200 with a nonzero base_resp.status_code. Check both layers. The official error-code reference should remain the source of truth for operational handling.

Production controls worth adding

  • Require the uploader to confirm recording, composition, lyric, performer, and voice permissions.
  • Store a rights record separately from the generated audio.
  • Scan file type and size before creating a signed source URL.
  • Give extracted lyrics a visible human-approval step.
  • Record prompt and lyric revisions without storing the API key.
  • Expire feature IDs in your database before MiniMax’s 24-hour limit.
  • Rate-limit by user as well as by MiniMax account.
  • Download temporary output URLs into storage you control.
  • Provide a deletion path for source audio, extracted lyrics, and output.
  • Run a documented listening and rights-review test before making quality or commercial-use claims.

MiniMax Music Cover API FAQ

Can a new user start a Music Cover API integration?

No. MiniMax closed new paid Music Generation access on August 20, 2026 and discontinued music-cover-free. Existing paying users may continue if their account remains entitled.

Does the cover API use Music 3.0?

No. An eligible existing paying account uses music-cover; music-cover-free is discontinued. music-3.0 is text-to-music and the open-weight Music 3 checkpoint is not the hosted Cover workflow.

Can I change the lyrics in one step?

The Music Generation schema allows a 10–1,000-character lyrics value with a cover request. The two-step route is safer when changes matter because it exposes the ASR transcription and structure before generation.

Can I upload the reference with the File Management API?

The cover schemas document audio_url and audio_base64, not a File Management file_id. Do not add an undocumented file-ID workflow.

How long can the reference recording be?

MiniMax documents 6 seconds through 6 minutes and a maximum file size of 50 MB.

Does preprocessing create a permanent reusable asset?

No. cover_feature_id is valid for 24 hours. Store the authorized source and review state according to your own retention policy if repeat generation is required.

Official references and related guides