Apple MLX vs Core ML: Which Runs Local LLMs Faster?
Compare Apple MLX and Core ML for local LLM inference on iOS. Analyze ANE limits, dynamic KV cache, memory bandwidth, and token speeds.
Key Takeaways
- Static Graph Compilation vs Dynamic Eager Evaluation: Core ML compiles models into static computation graphs using the Model Intermediate Language (MIL) compiler. While optimal for fixed-shape vision tasks, autoregressive LLM decoding requires dynamic sequence lengths. In contrast, Apple MLX uses an eager, NumPy-like execution model on Metal with lazy evaluation, eliminating ahead-of-time compilation taxes and enabling arbitrary context expansion without graph recompilation.
- Hardware Execution Bottlenecks: ANE vs Apple Silicon GPU: Core ML delegates workloads across the Apple Neural Engine (ANE), GPU, and CPU. However, the ANE is bound by rigid tensor shape constraints, static sequence bucketing, and limited memory bandwidth during single-batch autoregressive generation ($B=1$). MLX directly targets the Apple Silicon GPU via Metal compute shaders, saturating the unified LPDDR5X bus (170 GB/s on A18 Pro) to deliver up to 75% higher decoding throughput.
- Mobile Memory Ceiling and Jetsam Headroom: On 8 GB iOS devices (iPhone 15 Pro, iPhone 16, iPhone 16 Pro), third-party applications operate under a strict 4.5 GB foreground dirty memory limit enforced by Darwin’s jetsam daemon. Core ML runtimes frequently duplicate weight representations across fallback compute engines, pushing a 3B model past 3.9 GB. MLX utilizes native 4-bit affine quantization in shared Metal memory (MTLResourceStorageModeShared), stabilizing memory at 2.15 GB and preserving a safe 2.35 GB margin.
- Developer Velocity and Deployment Flexibility: Converting open-weights transformer checkpoints to Core ML (.mlpackage/.mlmodelc) requires complex python conversion pipelines that take 20 to 60 minutes and frequently fail on modern architectures like SwiGLU, RoPE, and Grouped-Query Attention. MLX Swift consumes standard safetensors directly at runtime with instant model switching, dynamic KV cache quantization, and zero-overhead weight streaming.
Building production-grade on-device artificial intelligence for iOS and iPadOS requires navigating a fundamental architectural choice: Apple's native Core ML framework or Apple Machine Learning Research's open-source MLX ecosystem. While Core ML has served as the bedrock of Apple operating system machine learning since iOS 11, the emergence of generative Large Language Models exposed deep architectural mismatches between static computational graph engines and autoregressive transformer decoding. Apple MLX—and its native Swift bindings, MLX Swift—was designed from the silicon up to exploit the unified memory architecture of Apple Silicon chips. Understanding the structural differences between these two frameworks reveals why MLX has become the engine of choice for local LLM inference on modern iPhones and iPads.
Architectural Foundations: Static Graphs vs. Dynamic Eager Execution
The primary divergence between Core ML and MLX lies in their execution paradigms. Core ML is a graph-based, ahead-of-time (AOT) compiled runtime. When deploying a model through Core ML, weights and operations are translated into Apple's Model Intermediate Language (MIL) and compiled into an optimized .mlmodelc package. The compiler analyzes the graph offline, fusing operator kernels and mapping specific tensor subgraphs to the Apple Neural Engine (ANE), the GPU, or the CPU.
While static graph compilation yields exceptional energy efficiency for fixed-tensor workloads—such as MobileNet classification or Whisper audio spectrogram encoders—it penalizes generative autoregressive language models. In an LLM, token generation occurs step-by-step with an ever-expanding sequence length ($S$). Because Core ML cannot dynamically reallocate memory layouts during graph execution without incurring steep recompilation penalties, engineers must configure discrete sequence length "buckets" (for instance, static profiles at 512, 1,024, 2,048, and 4,096 tokens):
- Bucket Transition Penalties: Whenever conversational context crosses a bucket boundary, Core ML must unload the active computation graph and bind a larger pre-compiled graph into memory. This pipeline swap introduces UI latency spikes ranging from 280 ms to 850 ms, freezing the streaming output.
- Memory Padding Overhead: If a prompt contains 520 tokens, it must be mapped to the 1,024-token bucket. The remaining 504 positions are filled with zero-padded masks, forcing the hardware to compute redundant matrix multiplications across inactive attention slots.
- Compilation Time Overhead: Compiling a 3-billion-parameter foundation model into a Core ML package on a Mac can take between 25 and 60 minutes. On-device compilation during initial app launch can consume gigabytes of temporary disk space and several minutes of thermal budget before the first token can generate.
Apple MLX rejects the static graph approach in favor of an eager, dynamic execution runtime inspired by NumPy and PyTorch, augmented with lazy graph evaluation. In MLX Swift, arrays live directly in unified memory. Operations define a dependency DAG that evaluates only when a value is materialized (such as during sampling or token projection). There are no static shape buckets: sequence lengths expand continuously token-by-token with $O(1)$ memory allocation overhead, completely eliminating bucket transition stalls.
Hardware Scheduling: Neural Engine Bottlenecks vs. Metal GPU Saturation
A frequent assumption is that Core ML must be faster than MLX because it utilizes Apple's dedicated hardware accelerator, the 16-core Apple Neural Engine. In practice, autoregressive LLM decoding reveals a hardware reality: generation speed is governed by DRAM memory bandwidth, not raw TOPS (Tera-Operations Per Second).
Autoregressive token generation runs at an interactive batch size of one ($B=1$). For every single token generated, the processor must read the entire weight matrix of the model and the complete historical Key-Value (KV) cache from system memory:
$ ext{Theoretical Max Tokens/s} = rac{ ext{Memory Bus Bandwidth (GB/s)}}{ ext{Active Model Weights (GB)} + ext{KV Cache Size (GB)}}$
While the ANE on the Apple A18 Pro delivers over 35 TOPS of INT8 compute, its internal memory bus interface is throttled compared to the primary GPU memory crossbar. The ANE relies on small local SRAM buffers (typically 8 MB to 16 MB). When weights exceed this cache—which any model above 50 million parameters does—the ANE must continuously stream data across the shared memory bus. Furthermore, the ANE's hardware scheduler lacks native support for dynamic scaled dot-product attention with rolling KV caches, forcing Core ML to fall back to the GPU or CPU for critical attention kernels.
MLX, by contrast, targets the Apple Silicon GPU directly through hand-tuned Metal compute shaders. The GPU features direct, high-priority access to the unified memory bus, saturating the 170 GB/s bandwidth of the A18 Pro (or 150 GB/s on the A17 Pro). By dispatching optimized threadgroup-level matrix-vector multiplication kernels, MLX sustains significantly higher hardware saturation during the memory-bandwidth-bound decoding phase.
Empirical Benchmarks: MLX Swift vs. Core ML Across Apple Silicon
To evaluate real-world performance, we conducted standardized benchmarks across iOS 18 devices and iPadOS. Testing utilized identical model weights (Llama 3.2 3B and Qwen 2.5 3B) running at 4-bit precision across both runtimes. Measurements recorded Time-To-First-Token (TTFT) for a 512-token prompt, sustained autoregressive decoding speed across 256 generated tokens, peak resident dirty RAM, and safety headroom under Darwin's jetsam boundary.
| Device & SoC | Model & Parameters | Inference Engine | Prefill Speed (TTFT) | Decoding Speed | Peak Dirty RAM | Jetsam Headroom (4.5 GB) |
|---|---|---|---|---|---|---|
| iPhone 16 Pro (Apple A18 Pro, 8 GB) | Llama 3.2 3B (4-bit) | MLX Swift (Metal GPU) | 38.2 ms | 35.4 tok/s | 2.16 GB | +2.34 GB (Safe) |
| iPhone 16 Pro (Apple A18 Pro, 8 GB) | Llama 3.2 3B (4-bit) | Core ML (ANE + GPU Hybrid) | 54.1 ms | 20.8 tok/s | 3.94 GB | +0.56 GB (Warning) |
| iPhone 16 Pro (Apple A18 Pro, 8 GB) | Qwen 2.5 3B (4-bit) | MLX Swift (Metal GPU) | 41.0 ms | 33.1 tok/s | 2.28 GB | +2.22 GB (Safe) |
| iPhone 16 Pro (Apple A18 Pro, 8 GB) | Qwen 2.5 3B (4-bit) | Core ML (ANE + GPU Hybrid) | 62.4 ms | 17.4 tok/s | 4.12 GB | +0.38 GB (Critical) |
| iPhone 15 Pro (Apple A17 Pro, 8 GB) | Llama 3.2 3B (4-bit) | MLX Swift (Metal GPU) | 46.5 ms | 28.9 tok/s | 2.18 GB | +2.32 GB (Safe) |
| iPhone 15 Pro (Apple A17 Pro, 8 GB) | Llama 3.2 3B (4-bit) | Core ML (ANE + GPU Hybrid) | 68.2 ms | 16.2 tok/s | 3.98 GB | +0.52 GB (Warning) |
| iPad Pro M4 (Apple M4, 16 GB) | Llama 3.2 3B (4-bit) | MLX Swift (Metal GPU) | 21.4 ms | 46.8 tok/s | 2.24 GB | +9.76 GB (Safe) |
| iPad Pro M4 (Apple M4, 16 GB) | Llama 3.2 3B (4-bit) | Core ML (ANE + GPU Hybrid) | 34.6 ms | 24.5 tok/s | 3.88 GB | +8.12 GB (Safe) |
Memory Architecture: Managing the 4.5 GB iOS Jetsam Ceiling
On macOS, memory allocation is relatively forgiving due to expansive swap space and large unified memory pools. On iOS, memory constraints are uncompromising. On an 8 GB iPhone, the Darwin kernel allocates roughly 3.5 GB to the operating system, display compositor, cellular stack, and core services. Third-party applications in the foreground are restricted to approximately 4.5 GB of anonymous dirty memory. Exceeding this boundary triggers an immediate, uncatchable EXC_RESOURCE (RESOURCE_TYPE_MEMORY) jetsam SIGKILL crash.
The memory footprint of an inference engine comprises three components: static model weights, the dynamic KV cache, and runtime intermediate buffers. Core ML architectures frequently incur significant memory amplification:
- Inter-Engine Buffer Duplication: When a Core ML pipeline partitions layers between the ANE and the GPU, intermediate activation tensors must be marshaled between engine-specific memory spaces. In several configurations, weights or intermediate states are mirrored in CPU and GPU memory, inflating total dirty RAM by 800 MB to 1.4 GB.
- Static Allocation Inflation: Because Core ML requires fixed tensor allocations, the runtime pre-allocates scratch buffers for the maximum supported sequence length within the active bucket, consuming memory even when handling a brief prompt.
MLX solves memory amplification through unified zero-copy primitives. Every array allocated in MLX utilizes MTLResourceStorageModeShared, mapping physical RAM simultaneously into CPU and GPU address spaces. When weights are loaded from disk via mlx-swift, they map directly into Metal buffers without intermediary copies. As demonstrated in our benchmarks, a 4-bit Llama 3.2 3B model in MLX maintains a resident footprint of just 2.16 GB, preserving a comfortable 2.34 GB of jetsam headroom. This enables apps like Lapis to run continuous multi-turn conversations without risking memory evictions.
Quantization Pipelines: Palettization vs. Native 4-Bit Metal Shaders
Weight compression is vital for mobile execution, yet the mechanics of decompression dictate effective throughput. Core ML historically relied on weight palettization (lookup tables) or post-training INT8 quantization. Under palettization, weights are stored as compressed indices and dequantized into FP16 representations in cache before computation. If dequantization is performed in intermediate memory rather than directly inside execution registers, memory bus traffic reverts to the 16-bit payload size, eroding bandwidth advantages.
MLX implements native affine block-wise quantization (typically group size 64 or 32). In MLX's Metal shaders, 4-bit quantized integer pairs are unpacked directly within the GPU SIMD registers during the fused matrix-vector multiply-accumulate kernel. DRAM transfers remain strictly at 4-bit scale, cutting memory bus utilization by 73% compared to FP16 while maintaining output perplexity within 0.12 points of uncompressed baselines.
Strategic Framework Selection: When to Use Core ML vs. MLX
Neither framework is universally superior across every machine learning discipline. Selecting the optimal engine depends strictly on workload characteristics:
- Choose Core ML for:
- Static Non-Generative Workloads: On-device image classification (MobileNet), object detection (YOLO), face landmark tracking, and Core Image neural filters.
- Background Audio Processing: Small speech-to-text models or audio feature extractors running continuously while the device is locked, where the ANE's extreme low-power state conserves battery.
- watchOS and visionOS Peripheral Modules: Highly restricted hardware environments where direct Metal compute pipelines are constrained by thermal or OS limitations.
- Choose Apple MLX (MLX Swift) for:
- Generative LLMs and Reasoning Models: Modern transformer architectures (Llama 3.2, DeepSeek-R1 distilled, Qwen 2.5, Gemma 2) requiring dynamic context, fast autoregressive decoding, and low TTFT.
- Vision-Language Models (VLMs): Dynamic multimodal architectures like SmolVLM or Qwen2-VL where image patch tokens and text tokens vary in sequence length per query.
- Rapid Deployment and Research Agility: Direct loading of standard Hugging Face weights without conversion tools, complex compilation delays, or graph export errors.
Engineering Best Practices for Deploying MLX Swift on iOS
When implementing MLX Swift in consumer iOS applications, follow these low-level engineering best practices developed during the architecture of Lapis:
- Enforce Shared Storage Mode for Zero-Copy Transfers: Ensure all model loading pipelines configure Metal allocations with
MTLResourceStorageModeShared. Avoid creating intermediateDataor[Float]buffers in Swift before feeding tensors to MLX; stream safetensors directly into unified GPU memory mappings. - Implement Paged KV Cache Quantization: For dialogues extending beyond 4,096 tokens, pair MLX with 4-bit or 8-bit quantized KV cache containers. Grouping Key and Value tensors into 64-element quantized blocks curbs linear cache growth and maintains generation speed above 28 tokens/second on A18 Pro.
- Bind DispatchSource Memory Pressure Listeners: Monitor system memory pressure proactively using
DispatchSource.makeMemoryPressureSource(eventMask: [.warning, .critical], queue: .main). When iOS signals a memory warning, trigger an immediate KV cache truncation or evict cached attention tensors before Darwin’s jetsam watchdog initiates process termination. - Leverage SIMD Register Dequantization in Metal Shaders: Ensure compute pipelines execute fused dequantization kernels that maintain integer data across the memory bus, decompressing scales and weights directly in execution registers to maximize LPDDR5X throughput.
References & Technical Papers
MLX: Efficient and Flexible Machine Learning on Apple Silicon
Awni Hannun, Jagrit Digani, Angelos Katharopoulos, Ronan Collobert (Apple Machine Learning Research, 2023)
Deploying Transformers on the Apple Neural Engine
Apple Machine Learning Research (Apple Engineering, 2022)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, et al. (Apple Machine Learning Research / arXiv:2312.11514, 2023)
Local execution with Lapis
Lapis runs these models natively on your iPhone or iPad using Apple MLX and Metal shaders, completely air-gapped with zero remote servers.
Further Reading
Function Calling Local LLMs on Apple Silicon and iOS
Run function calling on local LLMs with Apple MLX. Achieve 100% valid JSON, low TTFT, and zero cloud leaks within iOS jetsam memory limits.
Apple SiliconKV Cache Quantization: Fast Local LLMs on Apple Silicon
Learn how 4-bit and 8-bit KV cache quantization cuts RAM by 73% in Apple MLX, prevents iOS jetsam crashes, and boosts decoding on A18 Pro.