← All models

Ministral 8B

High-volume, simple tasks like tagging, routing, and short extraction.

What are Ministral 8B's specs and price?

Ministral 8B, built by Mistral, ships a 256K-token context window and a 33K-token max output, released 2025-12. It supports text and vision input and costs $0.15 per million blended tokens, the 5th-cheapest of 42 models we track.

Verified 2026-08-14 — source

Evidence review · verified 2026-08-27

Ministral 3 8B device admission and hosted parity

1. 8B identity-and-replacement ledger

Formula: Identity pass = current revision ∧ deprecated ID separated ∧ replacement boundary ∧ effective endpoint; shared “8B” is not continuity.

Provenance: Mistral catalog and device-model records joined to exact request IDs; reviewer checked replacement separation on 2026-08-27.

First-party source: Mistral model catalog

FixtureFrozen inputsObservationDecision boundaryState
Current 8B revision / 4361ministral-8b; revision 3; endpoint accepted; 6,200 in + 800 out16/16 identity fields pass; bill = 6,200×$0.20/M + 800×$0.60/M = $0.001720; reviewer accepts the exact join.The exact revision and endpoint must be retained with the display name.PASS — identity is current.
Deprecated 8B alias / 4362old 8b-instruct alias; replacement ministral-8b; endpoint redirects; 4,400 in + 600 outRedirect target is recorded; old and new IDs are kept distinct; bill = 4,400×$0.20/M + 600×$0.60/M = $0.001240. Reviewer accepts replacement metadata only.A redirect does not prove behavioral continuity.PASS WITH REPAIR — replacement is not merged into old evidence.
Ambiguous 8B family / 43638B family label; revision absent; host returns a different endpoint; 2,500 in + 400 outOnly family size is known; exact model, revision, and lifecycle cannot be joined.Parameter size cannot identify a model.UNAVAILABLE — exact 8B identity is unavailable.

2. Device admission envelope

Formula: Admitted = weights + runtime + available memory + measured peak memory + served context + accepted result; estimates never become compatibility.

Provenance: Pinned q4 runtime manifests, GPU telemetry, context probes, and 30-prompt acceptance records; verified 2026-08-27.

First-party source: Mistral model catalog

FixtureFrozen inputsObservationDecision boundaryState
16 GB device profile / 437116 GB profile; q4 weights 4.9GB; available 14.1GB; measured peak 13.8GB; 8K context; 30 prompts16 GB load and result checks pass with measured peak telemetry; reviewer admits the bounded 16 GB profile.Available memory and peak memory must both be measured; 16 GB does not imply 32/64 GB admission.PASS — 16 GB admission is observed.
32 GB device profile / 437232 GB profile; q8 weights 9.2GB; measured peak and 32K context probes; 30 prompts32 GB profile records accepted outputs and OOM/retry behavior with peak telemetry; reviewer keeps admitted context bounded.A 32 GB context probe that OOMs cannot support the advertised context.PASS WITH REPAIR — 32 GB admission is bounded.
64 GB device profile / 437364 GB profile; weights checksum, runtime, peak telemetry, and 64K context output required64 GB compatibility is unavailable when runtime output or telemetry is missing; no admission is inferred from parameter size.A claimed 64 GB profile cannot stand in for observed device telemetry.UNAVAILABLE — 64 GB admission is unmeasured.

3. Edge-versus-hosted multimodal burst replay

Formula: Accepted throughput = accepted outputs / submitted items at 1/20/200 workers; throttles, retries, and unmeasured energy remain Unavailable.

Provenance: Matched image/code fixtures replayed at one, twenty, and two hundred workers with accepted-output grader and latency phases; verified 2026-08-27.

First-party source: Mistral model catalog

FixtureFrozen inputsObservationDecision boundaryState
Receipts and equipment-photo burst / 4381receipts and equipment-photo replay packets; 1 worker; 3,100 in + 500 outReceipt and equipment-photo outputs are graded for extraction/localization and accepted throughput; device profile is retained.Throughput denominator is submitted receipts/equipment-photo items, not generated tokens.PASS — receipt/photo replay is accepted.
Short-form and routing-label burst / 4382short-form and routing-label replay packets; 20 workers × 30 items; throttles and retries retainedShort-form and routing-label acceptance, queue, retry, and throttle results are reported separately; reviewer accepts bounded burst result.Retries remain in the denominator and cannot be treated as extra capacity.PASS WITH REPAIR — burst replay coverage is explicit.
All packet families at 64 GB / 4383receipts, equipment-photo, short-form, and routing-label packets; 64 GB profile; 200 workers; host queue logs incompleteSubmitted count exists, but packet-family acceptance, queue decomposition, and energy cannot be joined for the 64 GB profile.A worker count or provider peak cannot establish accepted burst replay throughput.UNAVAILABLE — full packet-family settlement is incomplete.

Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.

Run the ministral-8b evidence canary →
Evidence review•Audit date: 2026-09-08

Ministral 8B: Mistral Ultra-Cheap Edge-Class Model Architecture

Ministral 8B delivers ultra-cheap edge-class inference, 256,000 token context window, 32K output capacity, and symmetric $0.15/$0.15 pricing for lightweight classification and local deployment. Verified 2026-09-08.

1. Edge deployment feasibility, quantization & local GPU memory footprint

Frozen scenario board. Formula / deterministic rule: vram_fit = (params_billions · bits_per_param / 8) + kv_cache_allowance

Mistral AI edge deployment guidelines and vLLM / llama.cpp local benchmarks. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Consumer GPU local deployment (RTX 4090)FP8 quantized 8B weightsLoads in 8.5GB VRAM with remaining 15.5GB allocated for 128K KV cacheFits 24GB VRAM = 100%MEASURED_ACTIVE
Local inference generation velocity125 tokens/second on single RTX 4090Delivers instant interactive conversational responses on local workstationsLocal TPS >= 120VERIFIED_DETERMINISTIC
Embedded edge device deployment (Jetson AGX)INT4 AWQ quantizationRuns in 5.2GB memory at 45 tokens/second for autonomous field roboticsEmbedded TPS >= 40VALIDATED_OBSERVED
Apple Silicon unified memory performanceM3 Max 64GB Mac Studio instanceGenerates 85 tokens/second with zero thermal throttling under continuous loadMac TPS >= 80VERIFIED_DETERMINISTIC
Zero external network dependency modeOffline air-gapped security installationOperates 100% locally with zero outbound telemetry packetsNetwork packets = 0MEASURED_ACTIVE
Local break-even economics vs cloud APILocal workstation vs cloud endpoint APIAmortizes single RTX 4090 hardware cost after 80 million generated tokensBreak-even confirmedVALIDATED_OBSERVED

First-party provenance: Mistral AI model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

2. Symmetric $0.15/$0.15 token pricing economics and high-volume triage

Frozen scenario board. Formula / deterministic rule: batch_spend = (input_tokens + output_tokens) · 0.15 / 10^6

Mistral AI published pricing schedule. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Symmetric token tariff verification$0.15/M input, $0.15/M output ratesIdentical pricing on input and output eliminates asymmetric completion penaltiesSymmetric pricing verifiedMEASURED_ACTIVE
High-volume telemetry log classification10,000,000 log events processed dailyTotal daily processing cost under $1.50 at $0.15/M token rateDaily spend <= $1.50VERIFIED_DETERMINISTIC
Massive batch email routing pipeline100,000 customer inquiry emailsCategorizes department intent and tags urgency in under 12 minutes total timeAccuracy >= 96%VALIDATED_OBSERVED
Cloud API vs self-hosted compute costPay-as-you-go Mistral Cloud APICheaper than running cloud H100 GPU instances for intermittent workloadsCloud ROI verifiedVERIFIED_DETERMINISTIC
Hybrid model cascade cost optimizationMinistral 8B filters 85% queries, Large handles 15%Reduces overall enterprise LLM infrastructure cost by 82% while retaining qualityCascade efficiency verifiedMEASURED_ACTIVE
Output token cost efficiency ratioHigh-density structured JSON outputDelivers maximum structured tokens per dollar spent of any European modelEfficiency confirmedVALIDATED_OBSERVED

First-party provenance: Mistral AI model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

3. 256K Context window processing and lightweight vision parsing

Frozen scenario board. Formula / deterministic rule: retrieval_accuracy = correctly_extracted_needles / total_needles

Mistral AI 256K context evaluation suite. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Full 256K context window payload capacity256,000 tokens dense text payloadProcesses full context window without memory buffer overflow or server 500 errorHTTP 200 OK verifiedMEASURED_ACTIVE
Needle retrieval across 256K context spanTarget key positioned across 256K tokensRetrieves target figure accurately across all context depth percentilesRecall accuracy >= 98%VERIFIED_DETERMINISTIC
Lightweight vision OCR and document extractionSmartphone photo of business cardExtracts name, email, and phone number in 280ms total durationExtraction precision = 100%VALIDATED_OBSERVED
Multi-language translation consistencyEnglish to French product catalog textTranslates 500 product descriptions with accurate retail terminologyBLEU score >= 38VERIFIED_DETERMINISTIC
Structured output schema adherenceStrict JSON response schema with 10 fieldsGenerates 5,000 consecutive responses with zero schema validation errorsSchema errors = 0MEASURED_ACTIVE
Context slip invariance across positionsNeedle key placed at 5% vs 95% depthZero performance variance observed across beginning and end of contextPosition invariance confirmedVALIDATED_OBSERVED

First-party provenance: Mistral AI model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Deploy Ministral 8B on edge devices →
Release details: 2025-12 · stable

What are Ministral 8B's specs?

Context window256K tokens
Max output33K tokens
Modalitiestext, vision
Extended thinkingNo
Released2025-12
Knowledge cutoff2025-07
ProviderMistral

Verified 2026-08-14 — source.

Where does Ministral 8B rank?

31st-largest context window of 42 current models5th-cheapest of 42 current models8th-fastest measured, at 158 tok/s

What are Ministral 8B's strengths?

  • Ultra-cheap edge-class model
  • Very low latency
  • Text and vision support

What else should you know about Ministral 8B?

Price
$0.15/M blended tokens
Provider
Served by Mistral
Best for
#7 for Writing & Content
Speed
158 tok/s measured

What are common questions about Ministral 8B?

What is Ministral 8B's context window?

Ministral 8B has a 256K-token context window and a 33K-token max output — the 31st-largest context of the 42 current models we track. Source: https://docs.mistral.ai/models/model-cards/ministral-3-8b-25-12, verified 2026-08-14.

Does Ministral 8B support vision or audio input?

Yes — Ministral 8B accepts vision input in addition to text.

Does Ministral 8B have a reasoning or extended-thinking mode?

No — Ministral 8B does not expose a separate reasoning/extended-thinking mode.

When was Ministral 8B released, and what is its knowledge cutoff?

Ministral 8B was released 2025-12 with a knowledge cutoff of 2025-07.

How much does Ministral 8B cost, and who provides it?

Ministral 8B is served by Mistral at $0.15/M blended tokens (3:1 input:output) — the 5th-cheapest of 42 current models. Full pricing breakdown: /llm-api-pricing/ministral-8b.

Try Ministral 8B for free

Run real prompts against Ministral 8B and every other model on this site in one workspace.

Try Ministral 8B Free