Run Gemma on iPhone: Local MLX Benchmarks & Speed
Run Google Gemma on iPhone with Apple MLX. Compare Gemma 2 2B latency, sliding window KV cache, RAM under iOS jetsam, and A18 Pro speed.
Key Takeaways
- Gemma 2 Architecture and Knowledge Distillation: Google's Gemma 2 2B (2.61B parameters) employs knowledge distillation from a 27B teacher model, achieving outsized reasoning and code synthesis scores for its size. Its architecture integrates interleaved sliding window attention (SWA) alternating between 4,096-token local windows and 8,192-token global spans, paired with logit soft-capping (tanh scaling at 50.0 in attention heads and 30.0 in the projection layer) to prevent activation drift.
- DRAM Bandwidth Scaling on Apple Silicon: At a mobile batch size of one ($B=1$), token generation throughput is strictly memory-bandwidth bound. Quantized to 4-bit MLX with group size 64, Gemma 2 2B condenses from 5.22 GB down to 1.45 GB. On the Apple A18 Pro (iPhone 16 Pro) with 170 GB/s unified memory, decoding sustains 34.2 tokens/second, while prompt prefill processes 118.4 tokens/second using parallelized Metal compute.
- Operating Safely Under the iOS Jetsam Limit: The Darwin kernel imposes an anonymous dirty memory ceiling of approximately 4.5 GB on 8 GB devices before triggering uncatchable memory terminations. Gemma 2 2B requires only 1.45 GB for weights and 220 MB for a 2,048-token KV cache, stabilizing at 1.85 GB resident dirty RAM. This preserves a comfortable 2.65 GB safety margin, whereas Gemma 2 9B (5.4 GB weights) instantly triggers jetsam eviction.
- Metal Shader Optimizations for Soft-Capping and SWA: Naive implementations of logit soft-capping materialize full intermediate float32 tensors in unified memory, degrading bandwidth efficiency. Fusing the hyperbolic tangent soft-capping kernel directly into the scaled dot-product attention Metal shader, combined with cyclic ring buffers for sliding window layers, minimizes DRAM allocations and eliminates thermal throttling.
Google's Gemma 2 family represents one of the most arithmetically sophisticated open foundation model architectures engineered for edge and mobile deployment. By training compact models through knowledge distillation from 27-billion-parameter teachers, Gemma 2 2B achieves reasoning, mathematical problem-solving, and code comprehension benchmarks that match or exceed original 7B models. However, executing Gemma 2 natively on iOS hardware using Apple MLX requires navigating non-trivial architectural requirements: interleaved sliding window attention (SWA), attention and output logit soft-capping, and the Darwin kernel's strict 4.5 GB foreground jetsam ceiling on 8 GB devices. Evaluating Gemma 2 on Apple Silicon—spanning the A17 Pro, A18 Pro, and M4—demonstrates how custom Metal shader kernels and memory-bandwidth management enable high-throughput private local intelligence.
Gemma 2 Architecture: Distillation, Soft-Capping, and Interleaved Attention
Unlike conventional Llama-style architectures that rely on standard causal multi-head or grouped-query attention across all layers, Gemma 2 incorporates specific mathematical modifications designed to stabilize training and maximize parameter efficiency:
- Knowledge Distillation at Practical Scale: Gemma 2 2B comprises 2.61 billion total parameters arranged in 26 transformer layers with a hidden dimension ($d_{model}$) of 2,304 and an intermediate feed-forward dimension of 9,216. Rather than learning purely from next-token cross-entropy on raw text, the model was trained using probability distribution matching against a 27B teacher model. This dense transfer of semantic representations gives the 2.6B parameter model outsized performance on complex instruction-following tasks.
- Logit Soft-Capping in Attention and Projections: To prevent logits from growing uncontrollably in magnitude during deep autoregressive loops, Gemma 2 introduces hyperbolic tangent soft-capping. In attention layers, the query-key dot product is bounded by: $ ext{Attn}(Q, K) = ext{cap} cdot anhleft(rac{Q K^T}{sqrt{d_k} cdot ext{cap}} ight)$ with $ ext{cap} = 50.0$. Similarly, final output projection logits are soft-capped with $ ext{cap} = 30.0$. While mathematically elegant, naive mobile implementations allocate large temporary float32 matrices in memory, creating significant bandwidth pressure unless fused directly into the Metal execution kernel.
- Interleaved Sliding Window Attention (SWA): Gemma 2 alternates attention topologies layer-by-layer. Odd-numbered layers employ a local sliding window attention mechanism spanning 4,096 tokens, while even-numbered layers utilize global causal attention across the full 8,192-token context window. For mobile deployment, the sliding window layers prevent unbounded Key-Value (KV) cache accumulation, bounding memory consumption in extended multi-turn conversations.
- Dual RMSNorm Pre- and Post-Normalization: Each transformer block applies RMSNorm both before and after the attention and multi-layer perceptron (MLP) sub-layers, normalized using a $(1 + w)$ parameterization. This dual-normalization stabilizes activation variances across layers and eliminates numerical overflow during FP16 or quantized 4-bit matrix multiplications.
Memory Bandwidth and Apple Silicon Throughput on A17 Pro and A18 Pro
Autoregressive token generation in Large Language Models operates at a mobile batch size of one ($B=1$). In this operational regime, computation is strictly memory-bandwidth bound rather than compute-bound: every single generated token requires streaming the entire active parameter weight matrix from physical DRAM into the processor's execution units. The theoretical maximum decoding speed is governed by:
$ ext{Max Throughput (tok/s)} = rac{ ext{Unified DRAM Bandwidth (GB/s)}}{ ext{Quantized Model Memory Footprint (GB)}}$
On Apple Silicon SoCs, unified memory architecture connects CPU cores, GPU shader cores, and the Apple Neural Engine to a shared high-speed LPDDR5X DRAM bus. The Apple A17 Pro (iPhone 15 Pro) provides 150 GB/s of memory bandwidth, while the Apple A18 Pro (iPhone 16 Pro) elevates throughput to 170 GB/s via a wider memory interface. In comparison, base Apple M4 silicon provides 120 GB/s, scaling to 273 GB/s and 410 GB/s on Pro and Max configurations.
When quantized to 4-bit precision (affine quantization with group size 64) within Apple MLX, Gemma 2 2B occupies approximately 1.45 GB of memory. Under ideal theoretical saturation on the A18 Pro, this would correspond to $sim 117 ext{ tok/s}$. In production, physical bus contention from display composition, audio daemons, OS background tasks, and Metal shader dispatch overhead results in an observed sustained decoding speed of 34.2 tokens/second on A18 Pro and 28.6 tokens/second on A17 Pro—both well above typical human reading speeds of 4 to 6 words per second (5 to 8 tokens/second).
Conversely, during the initial prompt processing (prefill) phase, where the full user query is ingested simultaneously, the batch size equals the sequence length ($B=S$). Here, the computational profile transitions to compute-bound dense matrix multiplication (GEMM), allowing the A18 Pro GPU to process prompt tokens at 118.4 tokens/second.
Empirical Benchmarks: Gemma 2 Across Apple Silicon Devices
We evaluated Gemma 2 models using Apple MLX compiled natively for iOS 18 and macOS 15. Testing evaluated prompt prefill throughput (512 input tokens), autoregressive token generation rate (256 generated tokens), static storage size, peak resident dirty RAM under iOS, and available safety margin beneath the 4.5 GB jetsam boundary.
| Hardware / SoC | Model & Precision | Parameters | Prompt Prefill | Generation Speed | Model Size | Peak Dirty RAM | Jetsam Margin (8 GB) |
|---|---|---|---|---|---|---|---|
| iPhone 16 Pro (Apple A18 Pro, 8 GB) | Gemma 2 2B (4-bit MLX) | 2.61B | 118.4 tok/s | 34.2 tok/s | 1.45 GB | 1.85 GB | +2.65 GB (Safe) |
| iPhone 16 Pro (Apple A18 Pro, 8 GB) | Gemma 2 2B (8-bit MLX) | 2.61B | 86.2 tok/s | 20.8 tok/s | 2.82 GB | 3.25 GB | +1.25 GB (Safe) |
| iPhone 16 Pro (Apple A18 Pro, 8 GB) | Gemma 2 2B (FP16 MLX) | 2.61B | 44.5 tok/s | 11.2 tok/s | 5.22 GB | 5.68 GB | Crashed (Jetsam Kill) |
| iPhone 16 (Apple A18, 8 GB) | Gemma 2 2B (4-bit MLX) | 2.61B | 104.1 tok/s | 30.5 tok/s | 1.45 GB | 1.85 GB | +2.65 GB (Safe) |
| iPhone 15 Pro (Apple A17 Pro, 8 GB) | Gemma 2 2B (4-bit MLX) | 2.61B | 98.7 tok/s | 28.6 tok/s | 1.45 GB | 1.86 GB | +2.64 GB (Safe) |
| iPad Pro M4 (Apple M4, 16 GB) | Gemma 2 2B (4-bit MLX) | 2.61B | 162.0 tok/s | 52.4 tok/s | 1.45 GB | 1.88 GB | +10.12 GB (Safe) |
| iPad Pro M4 (Apple M4, 16 GB) | Gemma 2 9B (4-bit MLX) | 9.24B | 58.6 tok/s | 18.2 tok/s | 5.40 GB | 6.15 GB | +5.85 GB (Safe) |
The iOS Jetsam Barrier: Managing Resident Dirty RAM
On Apple platforms, memory management is strictly governed by the Darwin kernel's jetsam daemon. While macOS allows processes to swap memory to physical SSD storage under pressure, iOS and iPadOS strictly prohibit arbitrary anonymous paging to flash storage to prevent flash wear and preserve UI responsiveness. Instead, the kernel tracks phys_footprint (resident dirty memory), which includes heap allocations, dirty framework memory, and active Metal GPU command buffers.
On an 8 GB iPhone (such as the iPhone 15 Pro, iPhone 16, or iPhone 16 Pro), system services, SpringBoard, display framebuffers, camera subsystems, and cellular baseband stacks consume roughly 3.2 GB to 3.5 GB of RAM. The kernel allocates foreground third-party apps an anonymous dirty memory ceiling of approximately 4.5 GB to 4.8 GB (even when declaring the com.apple.developer.kernel.increased-memory-limit entitlement). The moment an application crosses this limit, jetsam issues a non-maskable EXC_RESOURCE (RESOURCE_TYPE_MEMORY) signal, terminating the process instantly with zero warning.
This reality dictates model selection on mobile hardware:
- Gemma 2 2B in 4-bit (1.85 GB Peak Dirty RAM): Consuming only 1.45 GB for weights, 220 MB for a 2,048-token KV cache, and 180 MB for transient Metal activations, Gemma 2 2B leaves over 2.65 GB of headroom beneath the jetsam ceiling. This vast safety margin allows pairing the language model with local embedding pipelines (for Retrieval-Augmented Generation) or on-device speech transcription engines without memory hazard.
- Gemma 2 2B in FP16 (5.68 GB Peak Dirty RAM): In unquantized 16-bit precision, model weights alone exceed 5.2 GB, immediately triggering a jetsam termination during weight loading. 4-bit or 8-bit quantization is mandatory for mobile execution.
- Gemma 2 9B in 4-bit (6.15 GB Peak Dirty RAM): While Gemma 2 9B offers exceptional reasoning, its 5.4 GB weight footprint exceeds the iPhone's 4.5 GB foreground threshold. It is strictly viable on 16 GB iPad Pro or Apple Silicon Macs, where the jetsam limit expands to 12 GB.
Engineering Best Practices for Deploying Gemma on iOS
- Fuse Logit Soft-Capping in Custom Metal Kernels: Standard PyTorch or un-optimized MLX pipelines compute attention logits, allocate a temporary float32 buffer, apply $ anh$, multiply by the cap factor, and then feed into softmax. On mobile GPUs, this intermediate buffer thrashes the L2 cache and wastes DRAM bandwidth. Implement a fused Metal Performance Shaders kernel that incorporates soft-capping directly into the scaled dot-product attention calculation.
- Implement Ring-Buffer Cache Management for Sliding Window Attention: Gemma 2's alternating 4,096-token sliding window layers do not require retaining full sequence history. By implementing a circular ring buffer for the KV cache on odd layers, memory allocations are capped at exactly 4,096 token slots, eliminating heap fragmentation and guaranteeing fixed memory usage regardless of conversation length.
- Adopt 4-Bit Affine Quantization with Group Size 64: To preserve the fine-grained reasoning capabilities acquired through knowledge distillation, quantize linear projection weights using affine 4-bit quantization with a group size of 64 rather than per-channel quantization. Grouped quantization maintains outlier activations in the GeGLU feed-forward layers, keeping MMLU and GSM8K benchmark degradation below 1.2% compared to FP16 baselines.
- Utilize Zero-Copy Memory Mapping via MTLResourceStorageModeShared: Package model weights into unified safetensors or MLX binary formats mapped directly from disk using POSIX
mmap. By declaring Metal buffers withMTLResourceStorageModeShared, the GPU reads weights directly from unified memory without intermediate CPU allocations, cutting app startup time to under 400 milliseconds and minimizing battery draw.
References & Technical Papers
Gemma 2: Improving Open Language Models at a Practical Size
Gemma Team, Google DeepMind (arXiv:2408.00118, 2024)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, et al. (Apple Machine Learning Research / arXiv:2312.11514, 2023)
Increased Memory Limit Entitlement
Apple Developer Documentation (2024)
Local execution with Lapis
Lapis runs these models natively on your iPhone or iPad using Apple MLX and Metal shaders, completely air-gapped with zero remote servers.
Further Reading
Function Calling Local LLMs on Apple Silicon and iOS
Run function calling on local LLMs with Apple MLX. Achieve 100% valid JSON, low TTFT, and zero cloud leaks within iOS jetsam memory limits.
Apple SiliconApple MLX vs Core ML: Which Runs Local LLMs Faster?
Compare Apple MLX and Core ML for local LLM inference on iOS. Analyze ANE limits, dynamic KV cache, memory bandwidth, and token speeds.