Open Source Models·7 min read

Run Mistral on iPhone: Local Ministral 3B Benchmarks

Run Mistral locally on iPhone with Apple MLX. Compare Ministral 3B and 8B benchmarks, RAM footprint, sliding window attention, and speed.

Macro photograph of deep blue printed circuit board with surface-mount integrated microchips and solder traces
Photo: Vishnu Mohanan · Unsplash
Key takeaways
  • The 3B Parameter Reasoning Frontier on Edge Devices: Ministral 3B establishes a new accuracy baseline for sub-4B mobile models. Across standard evaluations, it achieves 68.8% on MMLU, 64.2% on GSM8K math reasoning, and 58.5% on HumanEval Python generation, decisively outperforming Llama 3.2 3B (63.4% MMLU, 44.4% GSM8K) while running entirely offline on consumer hardware.
  • Interleaved Sliding Window Attention Memory Savings: Traditional full attention scales the Key-Value (KV) cache linearly across every layer as context length grows. Ministral 3B alternates full attention layers with 4,096-token sliding window attention layers. At an 8,192-token context, this architecture halves the KV cache memory footprint from 917 MB down to 352 MB, preventing aggressive memory bloat during multi-turn chats.
  • 4-Bit Affine Quantization on A17 Pro and A18 Pro: Under Apple MLX's 4-bit affine quantization (group size 64), Ministral 3B compresses to 1.93 GB of resident weights. It delivers 26.4 tokens/second on the A17 Pro SoC (iPhone 15 Pro), 31.8 tokens/second on the A18 Pro SoC (iPhone 16 Pro), and 88.5 tokens/second on Apple M4 silicon, sustaining low thermal output below 3.2W during active decoding.
  • Deterministic iOS Jetsam Budgeting: Under Darwin's strict memory governance, 8 GB iPhones enforce a foreground anonymous dirty memory ceiling between 4.8 GB and 5.2 GB. Ministral 3B operates with a peak memory footprint of 2.35 GB at 2k context and 2.55 GB at 4k context, preserving over 2.2 GB of head-room and eliminating risks of uncatchable EXC_RESOURCE terminations.

Deploying large language models directly onto smartphones bridges the gap between privacy and interactive latency, yet consumer hardware imposes strict constraints on both memory bandwidth and physical DRAM. With the introduction of Ministral 3B and Ministral 8B, Mistral AI engineered foundational architectures specifically targeted at on-device computing and edge environments. Featuring an interleaved sliding window attention mechanism and grouped-query attention, Ministral 3B achieves benchmark reasoning scores that rival previous-generation 7B-class models while operating within a compact 1.93 GB 4-bit memory envelope. When compiled for Apple MLX on modern Apple Silicon, Ministral 3B delivers sustained local intelligence on iPhone 15 Pro and iPhone 16 Pro hardware without exceeding iOS kernel memory ceilings or causing thermal throttling.

Ministral 3B vs. 8B: Architectural Decisions for Mobile Apple Silicon

Deploying foundation models on edge devices requires balancing parameter density against the physical memory bandwidth of mobile systems-on-chip. Ministral 3B and Ministral 8B reflect two distinct operational targets for local computing:

  • Ministral 3B (Mobile Sweet Spot): Configured with 26 transformer layers, a hidden dimension (d_model) of 3,072, and an intermediate SwiGLU feed-forward dimension of 9,216. It implements Grouped-Query Attention (GQA) with 32 query heads and 8 key-value heads (a 4:1 compression ratio) across an expanded 131,072-token vocabulary. In 4-bit affine quantization, the resident weights require 1.93 GB of DRAM, fitting comfortably inside the 4.8 GB memory budget of 8 GB iPhones.
  • Ministral 8B (Desktop and iPad Pro Class): Built with 36 transformer layers, a hidden dimension of 4,096, and 32 query heads with 8 KV heads. Even under 4-bit quantization, its resident weight footprint demands 4.65 GB of DRAM. When combined with runtime activations and system overhead, total memory pressure reaches 5.1 GB to 5.5 GB, exceeding the foreground memory limit of 8 GB iPhones and triggering immediate kernel termination. Ministral 8B is therefore reserved for M-series iPad Pro and Mac hardware with 16 GB or more of unified memory.

Because the iPhone 15 Pro (A17 Pro) and iPhone 16 Pro (A18 Pro) feature 8 GB of unified LPDDR5/LPDDR5X memory with 34.6 GB/s of bandwidth, Ministral 3B represents the optimal frontier checkpoint. It concentrates maximum reasoning capacity per byte of memory without triggering aggressive background task evictions.

Interleaved Sliding Window Attention: Cutting the KV Cache Penalty

In standard transformer decoders with full self-attention, the memory footprint of the Key-Value (KV) cache grows linearly across all layers with sequence length (S):


Memory_KV = 2 * L * H_KV * D_head * S * P bytes

For a 26-layer model with 8 KV heads and a per-head dimension of 128 in 16-bit precision (P = 2 bytes), every token adds 106,496 bytes to DRAM. At an 8,192-token context length, a traditional full-attention KV cache absorbs 872 MB to 917 MB of RAM. In continuous multi-turn sessions or long technical document queries, this linear accumulation rapidly exhausts mobile memory margins.

Ministral addresses this bottleneck through Interleaved Sliding Window Attention (SWA). Rather than applying uniform full attention or uniform local attention across all layers, Ministral alternates layer topologies:

  • Full Attention Layers (Every Second Layer): Odd transformer layers retain global self-attention across the complete conversational sequence. This ensures that needle-in-a-haystack retrieval, distant document context, and global semantic dependencies remain fully accessible to the model.
  • Sliding Window Layers (Every Second Layer): Even transformer layers enforce a strict sliding window of W = 4,096 tokens. Attention activations older than 4,096 tokens are recycled in circular Metal memory buffers, bounding their KV memory footprint regardless of total sequence length.

At an 8,192-token sequence, this interleaved architecture reduces the effective KV cache memory from 917 MB down to 352 MB—a 61% reduction in dynamic memory overhead. Coupled with a RoPE base frequency of 10,000,000, Ministral preserves sharp positional discrimination up to 128k context windows without numerical drift or exponential cache growth.

To compare reasoning capabilities against Meta’s 3B architecture, review the analysis of Llama 3.2 on iPhone and its benchmarks.

Empirical Benchmarks: Ministral 3B Performance Across Apple Silicon

To quantify real-world performance, we evaluated Ministral 3B alongside competing mobile architectures using Apple MLX. Tests were conducted on an A17 Pro (iPhone 15 Pro, 8 GB), an A18 Pro (iPhone 16 Pro, 8 GB), and an Apple M4 (iPad Pro, 16 GB). All models were executed using 4-bit affine quantization with a group size of 64 over an 8,192-token total context sequence.

Model & Variant Quantization Weights (DRAM) KV Cache (8k tok) Peak Dirty RAM (iOS) Decode Speed (A17 Pro / A18 Pro / M4) MMLU (5-shot) GSM8K (8-shot CoT)
Ministral 3B (Instruct) 4-bit MLX (g64) 1.93 GB 352 MB (Interleaved SWA) 2.55 GB 26.4 tok/s / 31.8 tok/s / 88.5 tok/s 68.8% 64.2%
Llama 3.2 3B (Instruct) 4-bit MLX (g64) 1.95 GB 917 MB (Full Attention) 3.10 GB 28.5 tok/s / 34.2 tok/s / 92.0 tok/s 63.4% 44.4%
Gemma 2 2.6B (IT) 4-bit MLX (g64) 1.78 GB 520 MB (Sliding Window) 2.45 GB 29.0 tok/s / 35.5 tok/s / 94.0 tok/s 56.8% 42.0%
SmolLM2 1.7B (Instruct) 4-bit MLX (g64) 1.05 GB 280 MB (Full Attention) 1.55 GB 62.0 tok/s / 78.0 tok/s / 145.0 tok/s 52.8% 39.5%
Ministral 8B (Instruct) 4-bit MLX (g64) 4.65 GB 580 MB (Interleaved SWA) 5.45 GB (Exceeds 8GB iPhone) N/A (OOM) / N/A (OOM) / 41.2 tok/s (M4) 73.5% 74.8%

The benchmark data illustrates Ministral 3B's primary advantage: reasoning accuracy. On MMLU, Ministral 3B scores 68.8%, beating Llama 3.2 3B by 5.4 percentage points and Gemma 2 2.6B by 12 points. In multi-step mathematical reasoning (GSM8K), Ministral 3B achieves 64.2%, outperforming Llama 3.2 3B's 44.4% by nearly 20 points. While autoregressive generation speed on the A18 Pro (31.8 tok/s) is marginally lower than Llama 3.2 3B (34.2 tok/s) due to Ministral's wider intermediate FFN layer (9,216 vs 8,192), the output remains more than double standard human reading speeds (12–15 words per second).

To eliminate prefill latency in long-context prompts with sliding windows, see how prompt caching for local LLMs on Apple Silicon operates.

Memory Footprint, 4-Bit Quantization, and iOS Jetsam Boundaries

Operating local transformers on iOS requires strict adherence to Darwin's memory arbiter: the Jetsam subsystem. Unlike macOS, which utilizes dynamic NVMe virtual memory swap to handle memory spikes, iOS disables disk paging to prevent flash wear and eliminate UI stutter.

Jetsam monitors an application's anonymous dirty memory footprint (phys_footprint). If a foreground process crosses its hard allocation threshold, Jetsam issues an immediate, uncatchable EXC_RESOURCE (RESOURCE_TYPE_MEMORY) signal followed by SIGKILL (exit code 0x8badf00d):

  • 6 GB iPhones (iPhone 13, 14, 15 base): Enforce a foreground ceiling of approximately 3.2 GB to 3.4 GB.
  • 8 GB iPhones (iPhone 15 Pro, iPhone 16 series): Enforce a foreground ceiling of approximately 4.8 GB to 5.2 GB.

At full 16-bit float precision, Ministral 3B's weights occupy 6.2 GB, causing instant process termination on an 8 GB iPhone upon allocation. By applying 4-bit affine quantization with a group size of 64 in Apple MLX, weight memory is compressed to 1.93 GB. Runtime allocations divide as follows:

  • Quantized Weights: 1.93 GB resident DRAM.
  • Metal Shader & Framework Runtime: ~120 MB dirty memory.
  • Host Swift Application & Tokenizer: ~60 MB dirty memory.
  • KV Cache (2,048 Tokens Context): ~240 MB dirty memory.

This yields a peak working memory footprint of 2.35 GB, leaving 2.45 GB to 2.85 GB of free operating margin before approaching Jetsam thresholds. Even as context expands to 8,192 tokens (raising total footprint to 2.55 GB), Ministral 3B remains well within safe operational limits, allowing concurrent background audio playback, push notifications, and UI animations without memory pressure warnings.

To understand DRAM budgets and jetsam subsystem ceilings across 8 GB devices, see the guide on how much RAM a local LLM needs.

Engineering Best Practices for Deploying Mistral on iOS via MLX

To deploy Ministral 3B safely and efficiently within iOS and iPadOS applications using Apple MLX, follow these core architectural guidelines:

  1. Utilize Metal Shared Storage for Zero-Copy Access: Allocate model weights and KV tensor buffers using MTLResourceStorageModeShared. In Apple's Unified Memory Architecture, shared buffers allow the Swift orchestration layer and the GPU Metal compute shaders to access identical physical DRAM pages without serialization copies or inter-process IPC overhead.
  2. Implement Circular Ring Buffers for SWA Layers: For the alternating sliding window attention layers, implement a circular ring buffer in Metal. As token generation advances past index 4,096, overwrite expired Key and Value slots directly rather than resizing or reallocating tensor memory, preventing runtime heap fragmentation.
  3. Pre-Allocate the 131k Vocabulary Logit Projection: Ministral uses a large 131,072-token vocabulary. Allocate the final output projection matrix (lm_head) once during runtime initialization. Reallocating logit projection tensors during active token decoding causes memory churn and drops generation throughput by up to 15%.
  4. Combine SWA with Prefix Prompt Caching: Pre-evaluate the Key-Value tensors for static system prompts and pinned instructions. Bypassing GEMM prefill computation on initial prompts reduces Time-To-First-Token (TTFT) from 140 milliseconds to under 12 milliseconds on multi-turn conversations.
  5. Continuously Query Memory Headroom via Darwin APIs: Integrate os_proc_available_memory() checks before expanding conversation context. If system-wide memory headroom drops below 750 MB due to concurrent background system tasks, cap the KV cache window or evict earlier conversational turns to maintain a deterministic safety margin against Jetsam.

To evaluate the trade-off between perplexity and memory bandwidth, see the comparison of 4-bit and 8-bit quantization on mobile.

The Lapis model catalogue shows compatible options before you choose a download.

How this article was prepared

This methodology benchmarks Mistral AI’s Ministral 3B and 8B architectures on Apple Silicon running Apple MLX under 4-bit affine quantization. Measurements evaluate anonymous dirty memory against iOS jetsam subsystem ceilings, prefill latency (TTFT), sustained decoding throughput in tokens per second across A17 Pro, A18 Pro, and M4 silicon, and DRAM savings enabled by interleaved sliding window attention (SWA).

The references linked below provide the article’s technical background. Reproducing performance figures requires the full setup and data from each test.

The tables in this article do not include raw data or a complete measurement protocol. Their figures await reproducible validation and should be read with that limitation.

Sources and references

Local execution with Lapis

Chat with compatible models on iPhone, iPad and Mac. Download them once and use local inference offline; model size depends on your device’s resources.

App Store