MiniMax AI vs Google Gemini: Models, APIs, Pricing, and Best Fit

Last verified: August 26, 2026.

Current-model update: Google made gemini-3.7-flash generally available on August 13, 2026 and now presents it as the current Flash model for coding and agentic work. Gemini 3.6 Flash remains generally available, but it is no longer the correct lead model for a current MiniMax-versus-Gemini Flash comparison. Google’s discounted 3.7 Flash rates run through December 31, 2026; higher standard rates begin January 1, 2027 unless Google changes the schedule.

MiniMax AI vs Google Gemini: the short answer

For a current API comparison, test MiniMax M3 against Gemini 3.7 Flash. Both accept text, image, and video input and return text; Gemini 3.7 Flash additionally documents audio and PDF input. MiniMax M3 has lower listed direct input and output token rates in the compared bands during Google’s current promotion, while Gemini offers a broad Google developer ecosystem and a clearly documented 1,048,576-token input limit. Quality, latency, cache behavior, tool use, and total accepted-task cost still require a matched benchmark.

Do not merge Google surfaces. The model ID, token limits, and API prices below describe the Gemini API. They do not prove that the consumer Gemini app exposes the same selectable model, context, quota, or billing unit to every account.

QuestionMiniMax M3Gemini 3.7 Flash
Current hosted model to testMiniMax-M3gemini-3.7-flash
Release positionCurrent M-series hosted model in MiniMax’s catalogGenerally available since August 13, 2026; current Flash recommendation in Google’s model docs
Input typesText, image, and videoText, image, video, audio, and PDF
Output typeTextText
Published contextUp to one million combined input/output tokens; MiniMax also documents a guaranteed API minimum of 512K1,048,576 input tokens and up to 65,536 output tokens
Current listed input/output priceStandard: $0.30/$1.20 per 1M tokens at input ≤512K; $0.60/$2.40 above 512KPromotional through December 31: $0.75/$3.75 per 1M tokens
WeightsM3 weights under the MiniMax Community LicenseHosted Gemini API model; no downloadable 3.7 Flash weights established by the reviewed Google model page
Media generationSeparate MiniMax speech, image, and video models; hosted Music has a new-user restriction3.7 Flash’s listed modalities establish understanding inputs and text output, not image, audio, or video generation by the same model

Gemini 3.7 Flash is current; Gemini 3.6 is historical context

Gemini 3.6 Flash did not become false when 3.7 launched. It remains a generally available API model and may remain pinned in existing applications. The editorial correction is narrower: use 3.7 Flash for “latest,” “current,” “recommended Flash,” current pricing analysis, and the main verdict. Keep 3.6 only in clearly dated launch history, migration notes, or a deliberate backwards-compatibility test.

ModelCurrent roleCorrect article treatment
Gemini 3.7 FlashCurrent GA Flash model for coding and agentic workloadsLead Google model in the comparison, pricing table, recommendations, and FAQ.
Gemini 3.6 FlashStill GA, but superseded as the current Flash leadKeep in dated history or explicit migration comparisons; do not call it Google’s newest Flash model.
Gemini 3.5 Flash-LiteSeparate efficiency-oriented modelKeep only where the article intentionally compares the Lite tier; do not substitute it for 3.7 Flash.
Official Google Gemini 3.7 Flash model card
Google’s official Gemini 3.7 Flash model page, captured August 26, 2026.

Context and multimodal inputs

Google’s Gemini 3.7 Flash model page lists a 1,048,576-token input limit and 65,536-token output limit. It accepts text, images, video, audio, and PDFs and returns text. MiniMax M3 documents text, image, and video input, text output, and up to one million combined input/output tokens. These limits use different definitions, so a single “1M vs 1M” cell is incomplete.

CapabilityMiniMax M3Gemini 3.7 FlashTest that matters
Long textUp to 1M combined input/output documented1,048,576 input; 65,536 outputRetrieval accuracy at multiple positions, not capacity alone.
Image understandingYesYesOCR, charts, spatial detail, and abstention on ambiguous content.
Video understandingYesYesTemporal questions, sampling behavior, duration, and cost.
Audio understandingNot established for M3 by the compared model specYesSpeech, music, speaker, language, and long-audio test cases.
PDF inputUse an application-level extraction or supported multimodal route as documentedExplicitly listedScanned pages, tables, citations, and layout retention.
Generated responseTextTextDo not describe input understanding as media generation.

Coding and agents: benchmark 3.7, not yesterday’s default

Google positions Gemini 3.7 Flash for coding and agentic tasks. MiniMax positions M3 for tool use, agentic reasoning, and long-context work. Those provider descriptions identify intended use; they do not determine an overall winner. Run both models on the same repository snapshot with the same tool permissions and acceptance tests.

  • Pin gemini-3.7-flash and MiniMax-M3; do not let “latest” aliases change mid-run.
  • Record input, cached input, output—including Gemini thinking tokens where billed—tool calls, retries, wall time, and accepted patches.
  • Measure reviewer time and regression rate, not only tokens per request.
  • Test structured outputs, function calls, repository navigation, and recovery after a failed tool action.
  • Use the exact provider’s compatibility layer; similar API shapes do not guarantee identical parameters.

Media generation is separate from Gemini 3.7 Flash input support

Gemini 3.7 Flash can analyze several media types but its documented output is text. Google’s separate generation models and products should be evaluated separately. MiniMax likewise uses dedicated speech, image, and video model families rather than ordinary M3 text output. MiniMax’s hosted Music/Lyrics APIs need a dated warning: paid access closed to new users on August 20, while eligible existing paying users may continue and Music 3.0 open weights remain a separate route.

API pricing: Google’s 3.7 Flash rate is promotional through 2026

The current Gemini figure must carry its end date. Google’s official pricing page lists discounted Gemini 3.7 Flash rates through December 31, 2026, followed by standard rates on January 1, 2027. Output pricing includes thinking tokens. Prices below are per one million tokens unless the storage row says otherwise.

Model / periodInputCached input/readOutputCache storage
MiniMax M3 Standard, input ≤512K$0.30$0.06$1.20Use current MiniMax cache documentation
MiniMax M3 Standard, input >512K$0.60$0.12$2.40Use current MiniMax cache documentation
Gemini 3.7 Flash promotion through Dec 31, 2026$0.75$0.075$3.75, including thinking tokens$0.50 per 1M tokens per hour
Gemini 3.7 Flash standard from Jan 1, 2027$1.50$0.15$7.50, including thinking tokens$1.00 per 1M tokens per hour

Current direct-rate conclusion: MiniMax M3’s listed Standard input and output rates are lower than Gemini 3.7 Flash’s promotional rates in both MiniMax input bands. Gemini’s promotional cached-input rate is lower than M3’s higher-band cache-read rate but higher than M3’s lower-band rate. This is not a quality or total-cost verdict. Audio/video processing, cache storage, thinking tokens, tools, retries, latency tiers, and accepted-output rate can reverse a headline-token comparison.

The former $1.50 input / $7.50 output row must not remain labelled as today’s Gemini 3.6 or 3.7 Flash price. It is the scheduled 3.7 Flash standard rate beginning January 1, 2027. Google’s promotion also applies to 3.6 Flash, but that does not make 3.6 the current lead model.

Which should you choose?

RequirementBetter first testReason
Lowest listed direct text-token rate in this comparisonMiniMax M3Its current Standard input/output rows are below Gemini 3.7 Flash’s promotional rows.
Audio or PDF input in the same language-model APIGemini 3.7 FlashBoth input types are explicitly documented on the current model page.
Google-native developer ecosystemGemini 3.7 FlashIt is Google’s current GA Flash model for coding and agents.
Downloadable weights from the compared familyMiniMax M3MiniMax publishes M3 weights under the MiniMax Community License; review its conditions.
Speech, image, or video generation through one vendor’s dedicated APIsMiniMax firstMiniMax documents separate generation families; confirm each endpoint and price.
New hosted music APINeither by assumptionMiniMax closed new paid hosted access; Gemini 3.7 Flash is a text-output model.
Long-context coding or agent workBenchmark bothCapacity figures do not measure retrieval, tool reliability, accepted patches, or total task cost.

Verdict

Gemini 3.7 Flash—not 3.6 Flash—is now the right Google model for the main current comparison. It has broader documented input modalities and a strong Google-native developer path. MiniMax M3 lists lower direct token rates, downloadable weights under a community license, and access to MiniMax’s separate generation families. Neither wins on specifications alone. Pin the exact models, apply Google’s December 31 promotion boundary, qualify MiniMax Music access, and choose from matched task results.

Frequently asked questions

What is the current Gemini Flash model?

gemini-3.7-flash is Google’s current GA Flash model in the reviewed API documentation. Gemini 3.6 Flash remains GA but should be treated as a prior-generation option, not the latest model.

What is Gemini 3.7 Flash’s context window?

Google lists up to 1,048,576 input tokens and 65,536 output tokens. That differs from MiniMax M3’s combined-context wording, so compare input and output limits separately.

How much does Gemini 3.7 Flash cost?

Through December 31, 2026, Google lists $0.75 per million input tokens, $0.075 per million cached-input tokens, and $3.75 per million output tokens including thinking tokens. On January 1, 2027, the listed standard rates become $1.50, $0.15, and $7.50 respectively unless Google changes the schedule. Cache storage is separately listed.

Does Gemini 3.7 Flash generate images, audio, or video?

Its model page lists text output. It can accept image, video, audio, and PDF inputs, but that is understanding, not media generation. Use Google’s separate generation models for generated media.

Is MiniMax M3 cheaper than Gemini 3.7 Flash?

Its listed Standard direct input and output token rates are lower than Google’s current promotional 3.7 Flash rates. That does not prove a lower total task cost; compare cache, thinking tokens, media billing, tools, retries, latency, and accepted outputs.

Official sources