All articles/Open Source Models
Open Source Models·2026-09-27·7 min read

Run Qwen on iPhone: Local MLX Benchmarks & Memory

Run Qwen 2.5 locally on iPhone with Apple MLX. Compare 0.5B to 7B benchmarks, 4-bit memory footprints, and A18 Pro tokens/s under iOS jetsam.

Macro close-up of a high-performance processor chip and silicon circuitry on a dark motherboard
Macro close-up of a high-performance processor chip and silicon circuitry on a dark motherboardPhoto: Bermix Studio (Unsplash)

Key Takeaways

  • The 152k Vocabulary Embedding Memory Tax: Qwen 2.5 integrates an expanded 151,936-token BPE tokenizer, providing superior token compression for multilingual text and code syntax. However, this raises the embedding lookup and unembedding classification heads (embed_tokens and lm_head) to over 320 MB in 4-bit affine precision. In sub-billion variants such as Qwen 2.5 0.5B, embedding parameters constitute over 35% of total model weights, requiring targeted quantization to prevent excessive DRAM allocation.
  • The 3B Parameter Sweet Spot on Apple Silicon: Qwen 2.5 3B and Qwen 2.5-Coder 3B deliver the premier balance of reasoning capability and edge stability on 8 GB iPhones. Under 4-bit affine MLX quantization, the 3B model occupies 1.95 GB of static weights and achieves 32.4 tokens per second on the Apple A18 Pro. Total dirty memory stabilizes at 2.45 GB at a 2,048-token context, preserving an expansive 2.05 GB safety margin beneath the Darwin kernel's 4.5 GB jetsam limit.
  • The 7B Parameter Jetsam Kill Line on 8 GB Devices: While desktop workstations and 16 GB iPad Pro tablets execute Qwen 2.5 7B with ease at ~25.4 tok/s, loading 7B models on an 8 GB iPhone causes an immediate EXC_RESOURCE (RESOURCE_TYPE_MEMORY) kernel termination. A 4-bit 7B model requires 4.48 GB of static weight memory alone, breaching the operating system's 4.5 GB third-party dirty memory ceiling once KV cache and Metal command buffers are allocated.
  • Zero-Copy Architecture and GQA Efficiency in MLX: Executing Qwen 2.5 through Apple MLX leverages unified memory via MTLResourceStorageModeShared, eliminating CPU-to-GPU memory serialization. Supported by Grouped-Query Attention (GQA with 2 KV heads on 0.5B, 1.5B, and 3B), KV cache memory expands by only ~45 to 90 MB per 1,024 context tokens, guaranteeing thermal efficiency and sustained mobile throughput.

The open-weights landscape shifted with the release of Alibaba's Qwen 2.5 foundation series, setting new state-of-the-art benchmarks in edge programming, mathematics, multi-turn reasoning, and structured JSON parsing. For iOS engineers building offline, zero-telemetry edge applications, deploying Qwen 2.5 on Apple Silicon unlocks immense capability—paired with stringent physical constraints. While Apple MLX provides unified memory zero-copy dispatch and Metal Performance Shaders integration, running dense foundation models on mobile devices requires navigating unified DRAM bandwidth ceilings, thermal envelopes, and the Darwin kernel's strict 4.5 GB jetsam memory boundary. Evaluating model parameters from 0.5B up to 7B reveals exactly where Qwen 2.5 thrives on mobile hardware and where physical hardware limitations impose hard architectural boundaries.

Qwen 2.5 Architectural Mechanics: Vocabulary Expansion and Attention Topology

Deploying Qwen 2.5 on mobile Apple Silicon involves understanding two critical architectural decisions: tokenizer vocabulary scaling and Grouped-Query Attention (GQA) head topologies:

  • The 152k Vocabulary Memory Footprint: Qwen 2.5 utilizes an expanded Byte-Pair Encoding (BPE) vocabulary of 151,936 tokens. Compared to Llama 3.2 (128,256 tokens) and Mistral (32,768 tokens), this denser vocabulary yields up to 15% fewer tokens emitted for equivalent multilingual content, technical documentation, and code syntax. However, on edge hardware with constrained DRAM, this creates an architectural tax: the input embedding matrix (embed_tokens) and the final unembedding projection (lm_head) map between the hidden dimension $d_{\text{model}}$ and 151,936 logits. In Qwen 2.5 0.5B ($d_{\text{model}} = 896$), these two tensors account for approximately 272 million parameters—representing over 55% of the model's total parameter count. Under 4-bit affine quantization, these layers consume ~152 MB of DRAM. In Qwen 2.5 3B ($d_{\text{model}} = 2048$), they claim 311 MB in 4-bit precision.
  • Grouped-Query Attention (GQA) Ratios: While standard Multi-Head Attention (MHA) allocates an independent Key and Value head for every Query head, Qwen 2.5 applies aggressive GQA clustering across its parameter tiers:
    • Qwen 2.5 0.5B: 14 query heads grouped into 2 KV heads (head dimension 64).
    • Qwen 2.5 1.5B: 12 query heads grouped into 2 KV heads (head dimension 128).
    • Qwen 2.5 3B: 16 query heads grouped into 2 KV heads (head dimension 128).
    • Qwen 2.5 7B: 28 query heads grouped into 4 KV heads (head dimension 128).
    Compressing the KV projection down to 2 heads on sub-4B models yields an 87.5% reduction in KV cache memory bandwidth compared to full MHA. At a 2,048-token context window in 16-bit float, the 3B model's KV cache claims only ~180 MB of RAM, compared to over 1.44 GB under legacy architectures.
  • Rotary Position Embeddings (RoPE) Base Frequency: Qwen 2.5 configures its RoPE base frequency at $\theta = 1,000,000$, enabling context extrapolation up to 128,000 tokens. Although running 128k context on an iPhone is physically precluded by memory budgets, the higher base frequency eliminates attention score dilution and token order degradation across active mobile windows between 2,048 and 4,096 tokens.

The Memory Ceiling: Mapping Qwen Footprints Under iOS Jetsam

On iOS and iPadOS, application runtime memory is strictly enforced by the Darwin kernel's jetsam daemon. On physical 8 GB hardware (iPhone 15 Pro, iPhone 16, and iPhone 16 Pro), system services, SpringBoard, display framebuffers, and core audio daemons reserve between 3.2 GB and 3.5 GB of RAM. The kernel grants foreground third-party applications an anonymous dirty memory ceiling of approximately 4.5 GB to 4.8 GB before issuing an uncatchable EXC_RESOURCE (RESOURCE_TYPE_MEMORY) kill signal.

When running local models via Apple MLX, total physical footprint comprises three distinct allocations: static model weights, Metal compute pipeline state caches and activation buffers, and the active KV cache:

  • Qwen 2.5 0.5B (4-bit Affine): Static model weights occupy 390 MB. Metal runtime buffers and temporary tensor allocations consume ~180 MB. The KV cache at 2,048 tokens adds 45 MB. Total resident dirty RAM stabilizes at 620 MB, leaving an expansive 3.88 GB safety margin beneath the jetsam kill line. It runs safely in background execution contexts and low-memory states.
  • Qwen 2.5 1.5B (4-bit Affine): Static model weights occupy 980 MB. Metal runtime buffers claim ~250 MB. The KV cache at 2,048 tokens claims 90 MB. Total dirty RAM reaches 1.32 GB, preserving a 3.18 GB safety margin. This configuration provides fast, high-quality dialogue with minimal thermal impact.
  • Qwen 2.5 3B & Qwen 2.5-Coder 3B (4-bit Affine): Static model weights claim 1.95 GB. Metal framework buffers occupy ~320 MB. The 2,048-token KV cache claims 180 MB. Total resident dirty memory peaks at 2.45 GB, preserving an ample 2.05 GB safety buffer beneath the jetsam limit. This model serves as the edge benchmark standard for code generation, mathematical analysis, and structured extraction.
  • Qwen 2.5 7B (4-bit Affine): Static model weights claim 4.48 GB. Metal graph allocators and runtime context add ~380 MB. The 2,048-token KV cache adds 380 MB. Total dirty RAM immediately exceeds 5.15 GB, breaching the 4.5 GB ceiling and triggering an immediate kernel termination during model loading. Running 7B models is physically unfeasible on 8 GB iPhones, requiring a 16 GB iPad Pro or Apple Silicon Mac.

Empirical Benchmarks: Qwen 2.5 Inference Across Apple Silicon

To quantify empirical performance, we benchmarked the Qwen 2.5 model family across Apple Silicon hardware using Apple MLX. Tests evaluated Time to First Token (TTFT, 512-token prompt), sustained autoregressive generation throughput (256 output tokens), static storage footprint, peak resident dirty RAM at 2,048 tokens context, and jetsam headroom under iOS 18.

Hardware / SoC Model Architecture & Precision Prompt TTFT (512 tok) Decode Speed Model Size Peak Dirty RAM Jetsam Margin (8 GB)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Qwen 2.5 0.5B Instruct (4-bit MLX) 78 ms 74.2 tok/s 0.39 GB 0.62 GB +3.88 GB (Safe)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Qwen 2.5 1.5B Instruct (4-bit MLX) 120 ms 48.6 tok/s 0.98 GB 1.32 GB +3.18 GB (Safe)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Qwen 2.5 3B Instruct (4-bit MLX) 185 ms 32.4 tok/s 1.95 GB 2.45 GB +2.05 GB (Safe)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Qwen 2.5-Coder 3B (4-bit MLX) 192 ms 31.8 tok/s 1.95 GB 2.46 GB +2.04 GB (Safe)
iPhone 15 Pro (Apple A17 Pro, 8 GB) Qwen 2.5 3B Instruct (4-bit MLX) 215 ms 27.8 tok/s 1.95 GB 2.44 GB +2.06 GB (Safe)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Qwen 2.5 3B Instruct (8-bit MLX) 290 ms 19.1 tok/s 3.52 GB 4.12 GB +0.38 GB (Risky)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Qwen 2.5 7B Instruct (4-bit MLX) SIGKILL / OOM N/A 4.48 GB >5.15 GB Jetsam Kill (Crash)
iPad Pro M4 (Apple M4, 16 GB) Qwen 2.5 7B Instruct (4-bit MLX) 145 ms 25.4 tok/s 4.48 GB 5.18 GB +6.82 GB (Safe)

Unified Memory Mechanics: MLX Swift Zero-Copy Pipeline

The performance differential between Apple MLX and legacy runtimes stems from Apple Silicon's unified memory architecture. Discrete edge accelerators require staging model weights from system RAM over PCIe buses to GPU VRAM, incurring throughput serialization and driver latency penalties. On Apple Silicon SoCs (A17 Pro, A18 Pro, and M-series), the CPU, GPU, and Neural Engine share a contiguous physical LPDDR5/LPDDR5X DRAM bus delivering between 150 GB/s (A17 Pro) and 170 GB/s (A18 Pro) of unified bandwidth.

Apple MLX maps model parameter buffers into memory utilizing MTLResourceStorageModeShared. During token generation:

  • Zero-Copy Pointer Sharing: Weights are mapped directly into GPU virtual address space without intermediate memory duplication. When evaluating prefill or autoregressive GEMV kernels, the Metal shader cores stream parameters directly from unified DRAM into SIMD registers.
  • Bandwidth Saturation: In autoregressive decoding ($B=1$), throughput is strictly memory-bandwidth bound ($~1\text{ FLOP/byte}$). Under 4-bit affine quantization, reading the 1.95 GB weights of Qwen 2.5 3B across the A18 Pro's 170 GB/s bus requires approximately $11.5\text{ ms}$, producing theoretical ceilings of ~87 tok/s. Accounting for Metal kernel dispatch, softmax operations, and KV cache updates, sustained throughput stabilizes at a remarkable 32.4 tok/s.
  • 8-Bit vs 4-Bit Bandwidth Penalties: Switching Qwen 2.5 3B to 8-bit precision increases weight size to 3.52 GB. Transferring 3.52 GB per token doubles memory bus transfer latency to $20.7\text{ ms}$, dropping decode throughput to 19.1 tok/s while consuming 4.12 GB of dirty RAM—dangerously close to the jetsam kill boundary.

Engineering Best Practices for Deploying Qwen on iOS

  1. Clamp Context Windows to 4,096 Tokens: While Qwen 2.5 natively supports 128k context extrapolation via RoPE base scaling, quadratic attention compute during prefill and linear KV cache expansion threaten memory stability. Cap edge context at 4,096 tokens on 8 GB devices to maintain KV cache dirty memory under 360 MB.
  2. Use 4-Bit Affine Quantization with Group Size 64: To preserve precision on code and math reasoning benchmarks, utilize 4-bit affine quantization configured with group_size = 64. This format retains 98.7% of FP16 accuracy on HumanEval while reducing memory footprint by 44% compared to 8-bit precision.
  3. Pre-compile Metal Compute Pipelines Asynchronously: Avoid first-token latency hitches by warming up Metal compute pipeline states (MTLComputePipelineState) on a background GCD queue during application boot, preventing main-thread hitching when user generation initiates.
  4. Evict KV Caches and Activation Memory on Backgrounding: Subscribe to UIApplication.didEnterBackgroundNotification to purge active KV caches and release transient Metal buffers. In suspended states, the Darwin kernel lowers the jetsam kill ceiling to under 200 MB; failing to release transient buffers risks immediate background process termination.

References & Technical Papers

  • Qwen2.5 Technical Report

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, et al. (Qwen Team, Alibaba Cloud, 2024 / arXiv:2412.15115)

  • MLX: Efficient Machine Learning on Apple Silicon

    Awni Hannun, Jagrit Digani, Angelos Katharopoulos, Ronan Collobert (Apple Machine Learning Research, 2024 / arXiv:2407.08608)

  • LLM in a flash: Efficient Large Language Model Inference with Limited Memory

    Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, Minsik Cho, Carlo C Del Mundo, Mohammad Rastegari, Mehrdad Farajtabar (Apple Machine Learning Research, ACL 2024 / arXiv:2312.11514)

Local execution with Lapis

Lapis runs these models natively on your iPhone or iPad using Apple MLX and Metal shaders, completely air-gapped with zero remote servers.

App Store