All articles/Apple Silicon
Apple Silicon·2026-09-26·7 min read

Speculative Decoding on Apple Silicon: MLX Speedup & Limits

Accelerate local LLM inference on Apple Silicon using MLX speculative decoding. Benchmark draft models, memory bandwidth, tokens/s, and iOS limits.

High-performance processor chip and micro-architecture mounted on a dark motherboard circuit board
High-performance processor chip and micro-architecture mounted on a dark motherboard circuit boardPhoto: Christian Wiediger (Unsplash)

Key Takeaways

  • Arithmetic Intensity Transformation in Unified Memory: Standard autoregressive decode is bounded by memory bus saturation (~1 FLOP/byte), reading gigabytes of model weights for each emitted token. Speculative decoding uses an ultra-compact draft model to emit candidate tokens serially, transforming target model verification into a parallel compute-bound GEMM operation ($O(\gamma)$ FLOPs/byte) executed in a single causal forward pass on the Apple Silicon GPU.
  • The Mobile Acceptance Rate (α) Threshold: While desktop GPUs with massive compute tolerate lower token acceptance rates, mobile Apple Silicon (A17 Pro and A18 Pro with 6 GPU cores) exhibits a sharp break-even boundary at $\alpha \approx 0.52$. Tasks with high structural predictability (code synthesis, JSON formatting) achieve $\alpha \ge 0.75$, delivering 1.6x to 1.8x speedups, whereas creative or reasoning tasks below the threshold suffer 10% to 18% throughput regressions due to wasted verification cycles.
  • Dual-Model DRAM Footprint vs 4.5 GB Jetsam Limit: Running speculative decoding on an 8 GB iPhone requires maintaining two resident models and two active KV caches. Pairing a 4-bit 3B target model (1.82 GB) with a 4-bit 135M draft model (95 MB) yields a peak dirty RAM footprint of 2.28 GB, leaving 2.22 GB of safety headroom beneath the Darwin kernel's 4.5 GB jetsam kill line. Conversely, a 1B draft model pushes memory to 2.85 GB, narrowing background safety margins.
  • Zero-Copy Pipeline Optimization in Apple MLX: Because Apple Silicon integrates unified memory across CPU, GPU, and Neural Engine without PCIe interconnects, MLX maps both draft and target models into shared DRAM buffers (MTLResourceStorageModeShared). Logit verification and KV cache updates occur in unified registers without host-device synchronization latency, maximizing mobile token generation rates.

When deploying generative Large Language Models (LLMs) on edge hardware, inference speed is fundamentally governed not by theoretical floating-point compute (TFLOPS), but by unified memory bandwidth. In standard autoregressive generation, emitting a single token requires streaming the model's entire parameter matrix across the memory bus from DRAM into processing registers. Speculative decoding addresses this physical bottleneck by decoupling inference into a fast drafting phase and a dense parallel verification phase. On Apple Silicon unified memory architectures, this technique introduces distinct hardware mechanics: while unified memory eliminates the PCIe serialization penalties common to discrete GPUs, mobile processors must balance constrained GPU core counts, thermal envelopes, and the Darwin kernel's strict 4.5 GB jetsam memory boundary. Understanding when speculative decoding accelerates edge inference—and when it induces net slowdowns—requires examining arithmetic intensity, token acceptance probability, and resident memory footprints.

The Memory Wall: Autoregressive Bottlenecks and Speculative Mechanics

Transformer text generation operates in two computationally distinct phases: prompt ingestion (prefill) and autoregressive token generation (decode). While prefill computes attention across hundreds of prompt tokens simultaneously in parallel matrix-matrix multiplications (GEMM), decoding executes sequentially with an active batch size of $B=1$:

  • Autoregressive Bandwidth Bottleneck: To produce each output token, the GPU must transfer all active model weights from physical DRAM to SIMD registers. For a 3-billion-parameter model quantized to 4-bit precision (approximately 1.82 GB), generating a single token requires reading 1.82 GB of data. The arithmetic intensity collapses to approximately 1 FLOP per byte transferred. On an iPhone 16 Pro equipped with an Apple A18 Pro delivering 170 GB/s of unified memory bandwidth, the physical ceiling on autoregressive throughput is $\text{Throughput} \le \frac{170\text{ GB/s}}{1.82\text{ GB}} \approx 93\text{ tok/s}$. In production runtimes with framework dispatch overhead and active attention calculations, sustained throughput stabilizes between 28 and 32 tokens per second.
  • The Speculative Drafting Phase: Speculative decoding (formalized by Leviathan et al. and Chen et al.) mitigates this constraint by introducing a tiny auxiliary model (the draft model, denoted $M_q$) alongside the primary target model ($M_p$). The draft model—typically 15x to 25x smaller than the target—generates $\gamma$ candidate tokens sequentially. Because an ultra-compact draft model such as SmolLM2 135M occupies only 95 MB in 4-bit precision, streaming its weights across the memory bus consumes negligible bandwidth, generating tokens at over 140 tok/s.
  • Parallel Verification: Rather than executing $\gamma$ sequential forward passes on the large target model, the target model ingests all $\gamma$ candidate tokens simultaneously in a single forward evaluation. Utilizing a causal attention mask, the target model computes the conditional probabilities $p(x_i \mid x_{

Verification follows modified rejection sampling to guarantee that the output token distribution matches the target model exactly without quality degradation. For each proposed token $x_i$, it is accepted with probability $\min\left(1, \frac{p(x_i)}{q(x_i)}\right)$. If a token is rejected at index $k$, subsequent candidate tokens are discarded, a replacement token is sampled from the adjusted distribution $\max(0, p(x) - q(x))$, and the loop repeats.

The Acceptance Rate Threshold: When Speculative Decoding Regresses on Mobile

The mathematical speedup factor of speculative decoding depends directly on the mean acceptance rate $\alpha \in [0, 1]$ and the draft window length $\gamma$:

$\mathbb{E}[\text{Accepted Tokens per Step}] = \frac{1 - \alpha^{\gamma + 1}}{1 - \alpha}$

On high-end desktop workstations equipped with discrete GPUs (such as an NVIDIA RTX 4090 or Apple Mac Studio M2 Ultra with 60+ GPU cores), raw compute capacity is abundant. Even if the draft model achieves only a mediocre acceptance rate ($\alpha \approx 0.45$), the target model's parallel verification step completes almost instantaneously, preserving net positive speedup.

On mobile Apple Silicon (iPhone and base iPad hardware), the performance dynamics shift fundamentally:

  • Thermal and Core Constraints: The Apple A17 Pro and A18 Pro incorporate 6 GPU cores operating within a strict mobile thermal envelope of 5 to 7 watts. While memory bandwidth is exceptionally high (150–170 GB/s), floating-point compute is not infinite. Generating $\gamma = 4$ candidate tokens with a draft model and running a 4-token parallel verification pass on the target model demands substantial ALU cycles.
  • The $\alpha = 0.52$ Break-Even Boundary: On an Apple A18 Pro running Llama 3.2 3B, drafting 4 tokens with a 1B model takes approximately 18 ms, while target verification requires 32 ms (totaling 50 ms for the step). If only 1 token is accepted ($\alpha \le 0.40$), the effective throughput drops to $\frac{1}{0.050\text{ s}} = 20\text{ tok/s}$—representing an 18% performance regression compared to standard single-model decoding at 28.6 tok/s.
  • Task Predictability Dictates Viability: Speculative decoding excels in structured domains where draft models align closely with target outputs: code synthesis ($\alpha \approx 0.76$), JSON data extraction ($\alpha \approx 0.81$), and structured entity extraction. Conversely, open-ended creative writing, complex mathematical reasoning, and non-English multilingual generation drop $\alpha$ below 0.45, inducing net latency penalties and higher battery drain.

Memory Footprint Under the 4.5 GB iOS Jetsam Ceiling

On iOS and iPadOS, application execution is strictly governed by the Darwin kernel's jetsam memory management daemon. On devices with 8 GB of physical RAM (iPhone 15 Pro, iPhone 16, and iPhone 16 Pro), system services, SpringBoard, and framebuffers claim approximately 3.2 GB to 3.5 GB of RAM. The kernel enforces an anonymous dirty memory ceiling of approximately 4.5 GB to 4.8 GB for foreground third-party applications.

Exceeding this boundary for even a single allocation cycle immediately triggers an untruncated EXC_RESOURCE (RESOURCE_TYPE_MEMORY) kernel kill signal. In speculative decoding, two distinct models and two independent Key-Value (KV) caches must reside in DRAM simultaneously:

  • Target Model Footprint: Llama 3.2 3B in 4-bit affine MLX format occupies 1.82 GB of static weight memory. Its active KV cache at a 2,048-token context adds approximately 330 MB, establishing a baseline footprint of 2.15 GB.
  • The Draft Model Choice (1B vs 135M):
    • Option A — Llama 3.2 1B (4-bit): Static weights claim 720 MB, and its 2,048-token KV cache claims 130 MB. Combined with target model allocations and Metal runtime overhead, peak resident dirty RAM reaches 2.85 GB to 3.05 GB. While functional, it leaves only 1.45 GB of headroom beneath the jetsam termination line, leaving the application vulnerable if memory spikes occur from background notifications or camera captures.
    • Option B — SmolLM2 135M (4-bit): Static weights occupy only 95 MB, and its KV cache claims 35 MB. Total resident dirty RAM remains tightly contained at 2.28 GB, preserving an expansive 2.22 GB safety margin beneath the jetsam kill threshold.

Because SmolLM2 135M maintains an architectural vocabulary compatible with modern Byte-Pair Encoding tokenizers and fits within the L2 cache structures of Apple Silicon during evaluation, ultra-compact draft models provide the optimal balance between high acceptance probability and bulletproof memory stability on mobile devices.

Empirical Benchmark: Speculative Decoding Speedup Across Apple Silicon

To quantify empirical performance, we benchmarked speculative decoding against standard autoregressive decoding across multiple Apple Silicon processors. Tests evaluated cold Time to First Token (TTFT, 500-token prompt), sustained generation speed (250 tokens emitted), mean acceptance rate ($\alpha$), peak resident dirty RAM, and jetsam safety headroom under production workloads.

Hardware / SoC Model Architecture & Precision Workload & Domain Acceptance Rate (α) Baseline Decode Speculative Decode Speedup Factor Peak Dirty RAM Jetsam Headroom (8 GB)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Target: Llama 3.2 3B (4-bit)
Draft: SmolLM2 135M (4-bit)
Python Code Completion / JSON Schema 0.76 28.6 tok/s 51.8 tok/s 1.81x 2.28 GB 2.22 GB (Safe)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Target: Llama 3.2 3B (4-bit)
Draft: Llama 3.2 1B (4-bit)
General Dialogue & Technical Q&A 0.64 28.6 tok/s 42.1 tok/s 1.47x 2.85 GB 1.65 GB (Safe)
iPhone 15 Pro (Apple A17 Pro, 8 GB) Target: Llama 3.2 3B (4-bit)
Draft: SmolLM2 135M (4-bit)
Technical Document Summarization 0.68 24.2 tok/s 40.7 tok/s 1.68x 2.26 GB 2.24 GB (Safe)
iPhone 15 Pro (Apple A17 Pro, 8 GB) Target: Llama 3.2 3B (4-bit)
Draft: Llama 3.2 1B (4-bit)
Abstract Creative Writing & Reasoning 0.42 24.2 tok/s 21.8 tok/s 0.90x (Regression) 2.84 GB 1.66 GB (Safe)
iPad Pro M4 (Apple M4, 16 GB) Target: Qwen 2.5 7B (4-bit)
Draft: Qwen 2.5 0.5B (4-bit)
Structured Tool Calling & SQL Querying 0.74 16.4 tok/s 31.2 tok/s 1.90x 4.85 GB 11.15 GB (Safe on 16 GB)
MacBook Pro (Apple M4 Pro, 24 GB) Target: Llama 3.3 70B (4-bit)
Draft: Llama 3.2 1B (4-bit)
Full Codebase Refactoring 0.78 7.1 tok/s 14.8 tok/s 2.08x 39.8 GB N/A (macOS VM)

Production Best Practices for Apple MLX Speculative Decoding

Deploying speculative decoding reliably in commercial iOS and iPadOS applications requires aligning execution parameters with the physical characteristics of Apple Silicon:

  1. Maintain a 15:1 to 25:1 Parameter Disparity Ratio: Selecting a draft model that is too large (such as a 1.5B draft for a 3B target) consumes excessive memory bandwidth during sequential drafting, eroding speculative gains. Selecting an ultra-compact draft model (such as 135M parameters for a 3B target, or 0.5B parameters for a 7B target) ensures drafting latency remains below 20% of the target verification step.
  2. Enforce Shared Tokenizer Vocabularies: Speculative verification assumes that draft tokens map directly to the target model's embedding space. If the draft and target models use distinct vocabularies or different Byte-Pair Encoding merges, token alignment fails, forcing expensive de-tokenization and re-tokenization cycles that eliminate performance benefits. Pair models from the same family (e.g., Llama 3.2 1B with Llama 3.2 3B, or Qwen 2.5 0.5B with Qwen 2.5 7B) or use distilled draft heads.
  3. Implement Adaptive Speculative Horizons (Dynamic $\gamma$): Static speculative horizons (such as fixing $\gamma = 5$) penalize performance during difficult sequence segments. Implement an adaptive controller that tracks the rolling acceptance rate over the preceding 10 tokens: expand $\gamma$ to 5 when $\alpha > 0.80$, reduce $\gamma$ to 2 or 3 when $\alpha$ drops between 0.55 and 0.70, and bypass speculative drafting entirely (reverting to pure autoregressive decoding) when $\alpha < 0.50$.
  4. Cap Combined Resident Dirty Memory at 3.0 GB: To guarantee crash-free background retention on 8 GB iPhones, configure MLX memory pools so that target weights, draft weights, and dual KV caches never exceed 3.0 GB of resident dirty RAM, preserving at least 1.5 GB of safety headroom beneath the jetsam kill boundary.

References & Technical Papers

  • Fast Inference from Transformers via Speculative Decoding

    Yaniv Leviathan, Matan Kalman, Yossi Matias (ICML 2023 / arXiv:2211.17192)

  • Accelerating Large Language Model Decoding with Speculative Sampling

    Charlie Chen, Sebastian Borgeaud, Alireza Ghaffarkhah, Samuel L. Smith (DeepMind, 2023 / arXiv:2302.01318)

  • MLX: Efficient Machine Learning on Apple Silicon

    Awni Hannun, Jagrit Digani, Angelos Katharopoulos, Ronan Collobert (Apple Machine Learning Research, 2024 / arXiv:2407.08608)

Local execution with Lapis

Lapis runs these models natively on your iPhone or iPad using Apple MLX and Metal shaders, completely air-gapped with zero remote servers.

App Store