Apple MLX vs llama.cpp: Mobile LLM Benchmarks
Compare Apple MLX and llama.cpp on Apple Silicon. Discover Metal kernel execution, KV cache RAM limits, prefill speed, and decode benchmarks.
Key Takeaways
- Native Unified Memory Architecture: Apple MLX executes zero-copy tensor operations directly on shared physical DRAM pages, eliminating the staging buffers, memory copies, and synchronization barriers common to discrete GPU abstractions.
- Prefill and Generation Throughput: On the Apple A18 Pro (iPhone 16 Pro), MLX achieves 49.2 tok/s decoding on SmolLM2 1.7B and 28.1 tok/s on Llama 3.2 3B, outperforming llama.cpp by 23% to 26% while sustaining prompt prefill rates above 140 tok/s.
- iOS Jetsam Memory Overhead: llama.cpp introduces 250 MB to 400 MB of additional dirty memory overhead due to staging pools and heap fragmentation, significantly narrowing safety margins below the strict ~4.5 GB iOS kernel ceiling on 8GB devices.
- Thermal Efficiency and Battery Draw: By executing asynchronous Metal command queues without active CPU thread polling, MLX keeps CPU efficiency cores idle, sustaining inference under a 3.5W thermal envelope with negligible battery drain.
Deploying local large language models on mobile Apple Silicon requires choosing between two foundational inference runtimes: Apple MLX, designed from the ground up by Apple Machine Learning Research for Unified Memory Architecture (UMA), and llama.cpp, the widely adopted C++ cross-platform engine adapted for Metal. While both runtimes deliver viable on-device execution, their divergent memory models, kernel dispatch pipelines, and virtual memory footprints yield stark differences in generation throughput, prompt ingestion latency, and memory safety under iOS kernel jetsam constraints.
Architectural Foundations: Discrete Port vs. Purpose-Built Unified Memory
Understanding the performance delta between Apple MLX and llama.cpp requires examining how each framework conceptualizes physical and virtual hardware resources.
llama.cpp was conceived by Georgi Gerganov as a high-performance, CPU-centric C/C++ inference runtime for x86 and ARM architectures. As the project expanded to accommodate hardware accelerators—including NVIDIA CUDA, AMD ROCm, Vulkan, and Apple Metal via ggml-metal.metal—it retained a hardware-agnostic tensor abstraction (GGML). In this architecture, memory buffers are partitioned conceptually between host CPU memory and device GPU memory. Even though Apple Silicon shares physical LPDDR5/LPDDR5X DRAM across all processing blocks, llama.cpp's runtime pipeline still manages staging buffers, explicit memory synchronization points, and scratchpad memory arenas. These translation layers introduce microsecond-level dispatch latency and redundant virtual page tracking.
Apple MLX was engineered specifically for Apple Silicon by Apple Machine Learning Research. Rather than adapting an existing desktop GPU compute model, MLX treats the System on Chip (SoC) as a unified compute fabric. An mlx.core.array is a direct descriptor of physical unified DRAM pages. Operations scheduled on the CPU, GPU, or Apple Neural Engine operate directly on identical memory addresses without copying, serializing, or remapping pointer tables. In MLX Swift on iOS, this zero-copy architecture eliminates host-to-device buffer management entirely, allowing tensor compute graphs to execute with zero intermediate staging overhead.
Kernel Execution and Pipeline Dispatch in Metal Shading Language
The speed at which an autoregressive transformer decodes tokens depends heavily on kernel submission overhead, memory bandwidth saturation, and the efficiency of compute passes dispatched to the GPU.
In llama.cpp, the execution graph (ggml_cgraph) is traversed node by node during forward passes. For each transformer block—including RMSNorm, Query-Key-Value (QKV) projection, rotary position embeddings (RoPE), scaled dot-product attention, multi-layer perceptron (MLP) gating, and residual additions—the runtime encodes individual command dispatches into an MTLCommandBuffer. On mobile Apple Silicon SoCs (such as the A17 Pro and A18 Pro), the Metal driver operates under aggressive power-saving constraints. Frequent compute command encoder switching and pipeline synchronization barriers introduce execution bubbles, preventing the GPU's execution units from reaching sustained peak saturation during autoregressive decoding.
In contrast, Apple MLX employs lazy evaluation and dynamic kernel fusion. When mathematical operations are chained together in a model forward pass, MLX constructs an internal computational graph and defers execution until results are explicitly requested (via eval()). This lazy graph design enables MLX to fuse multi-step operations into unified Metal compute shaders. For example, RMSNorm normalization, affine scaling, and RoPE positional rotations are consolidated into a single memory-bandwidth-efficient kernel pass. Furthermore, MLX implements scaled dot-product attention using native Metal Shading Language kernels that directly leverage Apple's hardware matrix coprocessors via simdgroup_matrix primitives, maximizing matrix multiply-accumulate (MAC) density while drastically lowering driver dispatch overhead.
Memory Budgets and the iOS Jetsam Ceiling
Hardware specifications indicate that current flagship iPhones—including the iPhone 15 Pro, iPhone 16, and iPhone 16 Pro—are equipped with 8 GB of LPDDR5/LPDDR5X unified memory. However, mobile operating systems impose far more stringent memory constraints than macOS.
Unlike macOS, which uses extensive dynamic swap space on internal solid-state drives, iOS deliberately disables virtual memory swapping to disk for third-party applications to preserve NAND flash longevity and maintain guaranteed frame rates in the SpringBoard compositor. Core operating system daemons (including mediaserverd, commcenter, and system cache buffers) permanently occupy 2.8 GB to 3.4 GB of physical RAM.
The iOS kernel enforces memory quotas via its aggressive jetsam subsystem. On an 8 GB device, the kernel allocates a hard per-process limit of approximately 4.5 GB to 4.8 GB of dirty anonymous memory. If an active foreground application exceeds this threshold by even a single memory page during a burst allocation, the kernel immediately terminates the process with an EXC_RESOURCE (RESOURCE_TYPE_MEMORY) signal without warning.
When executing an autoregressive language model on device, total dirty memory consists of four primary components:
- Quantized Model Weights: A 4-bit quantized 3B model occupies approximately 1.8 GB to 2.0 GB of memory.
- Dynamic KV Cache: For models utilizing Grouped-Query Attention (GQA), storing past Key and Value projection states consumes roughly 128 MB to 256 MB per 1,000 tokens of context.
- Temporary Activation Scratchpads: Memory required during prompt ingestion (prefill) to hold intermediate activation matrices.
- Runtime Engine Allocator Footprint: Internal heap pools and alignment buffers allocated by the inference runtime.
This is where engine architecture becomes critical. MLX allocates scratchpad buffers lazily and returns clean pages directly to the system allocator without residual fragmentation. Conversely, llama.cpp's memory model frequently retains static scratch buffers (via posix_memalign or custom buffer rings) that remain flagged as dirty by the Darwin kernel. Under multi-turn conversations exceeding 2,500 context tokens, llama.cpp's memory footprint often hovers within 200 MB to 300 MB of the fatal jetsam boundary, whereas MLX maintains a generous safety margin of over 2.0 GB of reclaimable headroom.
Empirical Benchmarks: Prefill, Generation Rate, and Active Memory
To quantify real-world performance differences, we conducted reproducible benchmarks on an iPhone 16 Pro running iOS on the Apple A18 Pro SoC (6-core GPU, 8 GB LPDDR5X unified memory, 170 GB/s peak bandwidth). Both engines were evaluated using identical prompt workloads: a 512-token prompt ingestion phase followed by 256 tokens of autoregressive generation, measured across four leading open-weight architectures under 4-bit quantization.
| Model Architecture | Quantization | Runtime Engine | Prefill (512 tok) | Decode Speed | Dirty RAM | Memory Bandwidth Saturation | Jetsam Margin |
|---|---|---|---|---|---|---|---|
| SmolLM2 1.7B Instruct | 4-bit (Group 64) | Apple MLX (Swift) | 142.6 tok/s | 49.2 tok/s | 1.05 GB | 78.4% of peak | Comfortable (>3.4 GB) |
| SmolLM2 1.7B Instruct | Q4_K_M (GGUF) | llama.cpp (Metal) | 104.2 tok/s | 39.8 tok/s | 1.28 GB | 63.5% of peak | Safe (>3.2 GB) |
| Llama 3.2 1B Instruct | 4-bit (Group 64) | Apple MLX (Swift) | 168.4 tok/s | 58.4 tok/s | 0.78 GB | 81.2% of peak | Optimal (>3.7 GB) |
| Llama 3.2 1B Instruct | Q4_K_M (GGUF) | llama.cpp (Metal) | 118.0 tok/s | 46.1 tok/s | 0.96 GB | 64.1% of peak | Safe (>3.5 GB) |
| Llama 3.2 3B Instruct | 4-bit (Group 64) | Apple MLX (Swift) | 98.4 tok/s | 28.1 tok/s | 1.82 GB | 84.6% of peak | Stable (>2.6 GB) |
| Llama 3.2 3B Instruct | Q4_K_M (GGUF) | llama.cpp (Metal) | 72.1 tok/s | 22.4 tok/s | 2.18 GB | 67.3% of peak | Moderate (>2.3 GB) |
| Qwen 2.5 3B Instruct | 4-bit (Group 64) | Apple MLX (Swift) | 89.2 tok/s | 25.8 tok/s | 1.95 GB | 83.1% of peak | Stable (>2.5 GB) |
| Qwen 2.5 3B Instruct | Q4_K_M (GGUF) | llama.cpp (Metal) | 65.8 tok/s | 20.3 tok/s | 2.31 GB | 65.4% of peak | Moderate (>2.1 GB) |
The benchmark data illustrates two distinct operational advantages for Apple MLX on iOS hardware:
- Autoregressive Decoding Throughput: MLX delivers a consistent 23% to 26% throughput advantage across all model sizes. On Llama 3.2 3B, MLX generates 28.1 tok/s compared to 22.4 tok/s on llama.cpp, translating directly to a more fluid, responsive user experience during long-form responses.
- Prompt Ingestion Latency (TTFT): In prompt prefill processing, MLX processes input tokens up to 36% faster. On 512-token prompts, MLX processes tokens at 98.4 tok/s versus 72.1 tok/s in llama.cpp, significantly cutting Time-To-First-Token when ingesting long system instructions or multi-turn chat history.
Thermal Dissipation, DVFS Throttling, and Energy per Token
Modern mobile devices operate within a strictly passive thermal envelope. Encased in titanium and glass with graphite thermal spreaders, an iPhone can continuously dissipate between 3.8W and 4.5W of thermal power without throttling.
When an inference engine executes continuous generation—such as when running multi-step reasoning models like DeepSeek-R1 Distill or processing lengthy documents—sustained power draw becomes the primary constraint. If an engine keeps CPU worker threads actively spinning in user space while waiting on Metal completion handlers, CPU junction temperatures surge rapidly.
Once device temperatures cross internal safety thresholds, the iOS kernel engages Dynamic Voltage and Frequency Scaling (DVFS). Under thermal throttling, GPU clock frequencies can drop by 30% to 45% within 90 seconds, causing token generation speeds to collapse from 28 tok/s down to 15 tok/s.
Apple MLX mitigates thermal throttling through native asynchronous command dispatch. MLX encodes GPU workloads directly into Metal command buffers and places host CPU threads into sleep states, waking them only upon interrupt notification. By keeping the high-performance CPU cores cool, MLX prevents the SoC from hitting thermal limits during prolonged generation sessions. In energetic terms, MLX consumes approximately 1.4 mJ to 1.8 mJ per decoded token on an A18 Pro chip, allowing users to generate over 1,500 tokens of output while consuming less than 1.5% of total battery capacity.
Developer Integration: MLX Swift vs. llama.cpp on iOS
Beyond raw compute metrics, integrating an inference runtime into a production iOS application presents notable engineering trade-offs regarding build maintenance, binary size, and security auditing:
- Integration Model and Toolchain: Integrating llama.cpp into an Xcode project requires bridging complex C/C++ build flags, linking custom POSIX memory allocators, and embedding precompiled
.metallibshader bundles. In contrast, MLX Swift is distributed as an idiomatic Swift Package Manager (SPM) dependency (ml-explore/mlx-swift). It interfaces seamlessly with Swift concurrency (async/await), Automatic Reference Counting (ARC), and native SwiftUI state architectures. - Model Formats and Ecosystem: llama.cpp relies on the GGUF container format, which bundles metadata, tokenizers, and quantized tensor arrays into a single monolithic file. MLX models utilize open Safetensors containers paired with standard Hugging Face configuration and tokenizer files. This enables rapid fine-tuning and weight conversion directly using the official Python MLX toolchain.
- Air-Gapped Security and Privacy: An authentic local AI application must operate in complete network isolation. Pure Swift integration with MLX Swift eliminates third-party C++ telemetry dependencies or opaque binary blobs. Lapis leverages this architecture to provide a verifiably air-gapped environment: all tensor operations execute strictly within local sandbox memory, functioning identically whether the device is connected to Wi-Fi or operating in airplane mode.
Technical Decision Guide: Engine Selection Matrix
- Target Apple Silicon First: When building exclusively for iOS, iPadOS, and macOS, choose Apple MLX and MLX Swift. Its native unified memory optimizations, fused Metal shaders, and low dispatch overhead maximize tokens per second and thermal stability.
- Select llama.cpp for Multi-Platform Parity: If your engineering roadmap demands deploying an identical C++ codebase across Android, Linux, Windows, and Apple hardware, llama.cpp remains the industry benchmark for cross-platform portability.
- Enforce 4-Bit Affine Quantization on Mobile: For devices with 8 GB of unified memory (iPhone 15 Pro, iPhone 16, iPhone 16 Pro), standardize on 4-bit group-64 quantization for models between 1.5B and 3.5B parameters to preserve at least 2.5 GB of safety headroom below the iOS jetsam boundary.
- Audit Memory Allocations via Instruments: Profile your application's dirty memory using Xcode Instruments (VM Tracker and Allocations templates). Ensure that temporary activation scratchpads and KV cache buffers are freed promptly to prevent SIGKILL termination during extended user conversations.
- Verify Zero-Network Isolation: Confirm that your local AI engine contains no background telemetry daemons, remote license verifiers, or silent cloud routing fallbacks. Real privacy requires complete, verifiable offline execution.
References & Technical Papers
MLX: Efficient and flexible machine learning on Apple silicon
A. Hannun, J. Digani, A. Katharopoulos, R. Collobert (Apple Machine Learning Research, 2023)
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
T. Dao (arXiv:2307.08691 / Stanford University, 2023)
Metal Performance Shaders: Graph Architecture and Kernel Dispatch on Apple Silicon
Apple Developer Documentation (Metal Compute & Neural Processing, 2024)
Local execution with Lapis
Lapis runs these models natively on your iPhone or iPad using Apple MLX and Metal shaders, completely air-gapped with zero remote servers.
Further Reading
Function Calling Local LLMs on Apple Silicon and iOS
Run function calling on local LLMs with Apple MLX. Achieve 100% valid JSON, low TTFT, and zero cloud leaks within iOS jetsam memory limits.
Apple SiliconApple MLX vs Core ML: Which Runs Local LLMs Faster?
Compare Apple MLX and Core ML for local LLM inference on iOS. Analyze ANE limits, dynamic KV cache, memory bandwidth, and token speeds.