All articles/Apple Silicon
Apple Silicon·2026-09-25·7 min read

Apple Neural Engine vs GPU for LLMs: Latency & Memory

Compare Apple Neural Engine vs GPU for local LLM inference. Analyze unified memory bandwidth, Core ML vs MLX, ANE constraints, and mobile tokens/s.

Turned-on MacBook Pro screen displaying low-level source code on a dark minimalist workstation desk
Turned-on MacBook Pro screen displaying low-level source code on a dark minimalist workstation deskPhoto: Arnold Francisca (Unsplash)

Key Takeaways

  • Arithmetic Intensity Dictates Hardware Placement: Autoregressive LLM generation operates at an arithmetic intensity near 1 FLOP/byte, making decoding strictly memory-bandwidth bound. While the Apple Neural Engine delivers 35 to 38 TOPS of peak INT8 matrix compute during parallel prefill, the Apple Silicon GPU's direct access to 150–170 GB/s of unified memory bandwidth delivers 3x to 4x higher sustained decode throughput.
  • Dynamic Context Versus Static Graph Compilation: Transformer KV caches expand incrementally with every generated token. Metal GPU pipelines handle arbitrary dynamic sequence lengths through zero-copy unified memory buffers (MTLBuffer). Conversely, the ANE compiler enforces static tensor shapes, requiring rigid context bucketing and zero-padding that waste memory and introduce multi-second latency spikes during bucket transitions.
  • Quantization Architecture Mismatch: Modern edge LLMs rely on 4-bit affine grouped quantization (group size 64/128) to preserve perplexity within tight mobile memory budgets. Metal kernels dequantize weights on the fly inside register threadgroups. The ANE lacks native affine 4-bit execution units, requiring intermediate unpacking passes that eliminate throughput gains.
  • The 4.5 GB iOS Jetsam Ceiling: On 8 GB iPhones, the Darwin kernel's jetsam daemon terminates foreground apps that exceed 4.5 GB of dirty memory. Core ML pipelines targeting the ANE often duplicate weights between execution runtimes, pushing 3B models past 4.2 GB. Native MLX implementations on Metal maintain a 2.15 GB dirty RAM footprint, preserving over 2.35 GB of headroom.

When Apple introduced modern Apple Silicon, marketing specifications prominently showcased the 16-core Apple Neural Engine (ANE), boasting compute throughput of 35 to 38 TOPS (trillion operations per second) on processors from the A17 Pro to the M4. For developers deploying edge intelligence, the natural assumption was to treat the ANE as the primary accelerator for running Large Language Models (LLMs). Yet in production runtimes such as Apple MLX, Ollama, and high-performance mobile frameworks, LLMs execute almost exclusively on the Apple Silicon GPU via Metal rather than the Neural Engine. This architectural reality is not an omission: it reflects the fundamental physics of transformer inference, rigid hardware constraints in the ANE datapath, and the Darwin kernel's strict memory governance on iOS.

Arithmetic Intensity: Why TOPS Do Not Translate to Generation Speed

Evaluating processor suitability for generative AI requires separating model execution into its two distinct physical phases: prompt ingestion (prefill) and autoregressive token generation (decode). Each phase operates with fundamentally different computational characteristics:

  • Prefill Phase (Compute-Bound GEMM): Ingesting an initial prompt involves parallel matrix multiplication across all input tokens simultaneously. The arithmetic intensity (ratio of floating-point operations to memory access) scales linearly with prompt length ($O(N)$). Here, high theoretical compute density—such as the ANE's 38 TOPS—delivers rapid tensor throughput when executing static matrix blocks.
  • Decode Phase (Memory-Bandwidth-Bound GEMV): Generating responses token by token is an autoregressive loop with an active batch size of $B=1$. To emit a single token, the system must stream the entire parameter weight matrix from DRAM into processing registers. The arithmetic intensity collapses to approximately 1 FLOP per byte transferred.

Because the decode phase dominates interactive user sessions, token generation throughput is strictly governed by memory bus saturation rather than raw arithmetic compute: $ ext{Throughput (tok/s)} approx rac{ ext{Memory Bandwidth (GB/s)}}{ ext{Model Weight Footprint (GB)}}$.

On Apple Silicon, the GPU is connected directly to the primary unified memory bus, commanding 150 GB/s on the A17 Pro, 170 GB/s on the A18 Pro, and between 120 GB/s and 546 GB/s across the M4 series. In contrast, the Apple Neural Engine interfaces with system DRAM through private DMA channels and a shared intermediate SRAM tile cache. Under sequential, single-token vector-matrix operations, the ANE suffers from memory pipeline starvation, yielding sustained generation speeds of only 8 to 14 tokens per second on a 3B model, whereas the GPU sustains 28 to 32 tokens per second on identical hardware.

The Static Shape Tax: Dynamic KV Caches vs Core ML Graph Compilation

Autoregressive transformer inference requires maintaining a Key-Value (KV) cache to avoid recalculating past token attention states. As dialogue progresses, this cache grows incrementally by one token position per forward pass.

The Apple Silicon GPU, programmed via Metal and Apple MLX, natively supports dynamic memory allocation. Paged or contiguous attention buffers (MTLBuffer) expand dynamically in unified RAM without requiring shader recompilation or pipeline re-initialization. Context lengths scale seamlessly from 1 token to 4,096 tokens with zero dispatch penalty.

The ANE, by contrast, relies on Apple's Model Intermediate Language (MIL) compiler (anecompiler), which requires strictly static tensor dimensions compiled ahead of time. To execute an LLM on the ANE via Core ML, developers must define fixed sequence length buckets (for example, discrete contexts of 512, 1,024, and 2,048 tokens). This rigid abstraction introduces severe operational overhead:

  • Zero-Padding Overhead: A prompt containing 520 tokens must be padded to 1,024 positions, wasting 49% of memory bandwidth and execution cycles computing dummy attention states.
  • Bucket Transition Latency Spikes: When a conversation extends beyond a bucket threshold, Core ML must unload the current execution plan and swap to a new static graph, inducing visible UI freezes of 250 to 900 milliseconds.
  • Cache Slicing Fragmentation: Managing static intermediate buffers on ANE introduces memory fragmentation that complicates multi-turn conversational retention.

Quantization Architecture Mismatch: 4-Bit Affine MLX vs Fixed-Precision ANE

Deploying multi-billion-parameter language models within the memory envelopes of mobile devices requires aggressive weight quantization. However, hardware support for quantization formats varies substantially across Apple Silicon compute blocks:

  • Native ANE Quantization (INT8 and FP16): The ANE's systolic execution units are optimized for dense INT8 and FP16 matrix operations with uniform, per-tensor or per-channel scaling factors. While Core ML supports compressed weight formats, non-native sub-8-bit weights must be unpacked into intermediate FP16 SRAM or DRAM buffers prior to execution, neutralizing memory bandwidth reductions.
  • GPU SIMDgroup Dynamic Dequantization (4-Bit Affine): In Apple MLX, models employ 4-bit affine grouped quantization (with group sizes of 64 or 128). Custom Metal Shading Language (MSL) kernels unpack 4-bit integer weights directly into GPU register threadgroups (simdgroup_matrix) on the fly during the GEMV execution loop. Weights remain in 4-bit format in physical DRAM, reducing memory footprint by over 70% without intermediate memory allocation.

This architectural disparity allows a 4-bit Llama 3.2 3B model in MLX to consume only 1.82 GB of static weight RAM, compared to 3.4 GB for an INT8 Core ML package targeting the ANE, while maintaining identical output perplexity within 0.12 points of the unquantized baseline.

Memory Footprint and the 4.5 GB iOS Jetsam Ceiling

On modern iPhones (iPhone 15 Pro, iPhone 16, and iPhone 16 Pro), physical DRAM is capped at 8 GB. The Darwin kernel strictly allocates system memory between display framebuffers, baseband firmware, SpringBoard, and system services, leaving third-party foreground applications governed by the jetsam memory daemon.

Jetsam enforces an anonymous dirty memory ceiling of approximately 4.5 GB to 4.8 GB. If resident dirty memory crosses this threshold for even a microsecond during prompt ingestion or KV cache allocation, the kernel immediately terminates the application process via an untruncated EXC_RESOURCE (RESOURCE_TYPE_MEMORY) signal.

When compiling and loading models via Core ML for ANE execution, the framework frequently creates duplicate weight representations across the CPU-GPU-ANE fallback runtime. A 3B-parameter model packaged in INT8 under Core ML routinely claims between 3.8 GB and 4.4 GB of dirty RAM. This leaves less than 700 MB of safety margin before triggering jetsam termination, making the application vulnerable to crashes whenever incoming notifications, background audio, or camera processes run concurrently.

By comparison, Apple MLX on Metal maps model weights directly from storage into unified memory using zero-copy memory mapping (mmap) with shared storage modes (MTLResourceStorageModeShared). Static weights occupy 1.82 GB, and active KV cache allocations at a 2,048-token context add 330 MB, establishing a peak resident dirty RAM footprint of 2.15 GB. The application maintains an extensive 2.35 GB buffer beneath the jetsam kill boundary, guaranteeing absolute operational stability.

Empirical Benchmark: Apple Neural Engine vs Metal GPU on Apple Silicon

To quantify empirical performance, we benchmarked identical transformer architectures across Core ML (targeting the 16-core ANE) and Apple MLX (targeting the Apple Silicon GPU via Metal). Tests measured cold Time-to-First-Token (500-token prompt), sustained generation speed (250 tokens generated), peak resident dirty memory, and jetsam headroom across multiple hardware tiers.

Hardware / SoC Runtime & Execution Target Model & Precision Time to First Token (TTFT) Sustained Generation Peak Dirty RAM Context Flexibility Jetsam Safety Headroom (8 GB)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Core ML (ANE Target) Llama 3.2 3B (INT8) 48 ms 10.4 tok/s 3.82 GB Static Buckets (512 / 1024 / 2048) 0.68 GB (High Risk)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Apple MLX (Metal GPU) Llama 3.2 3B (4-bit Affine) 42 ms 28.6 tok/s 2.15 GB Fully Dynamic ($O(1)$) 2.35 GB (Safe)
iPhone 15 Pro (Apple A17 Pro, 8 GB) Core ML (ANE Target) SmolLM2 1.7B (INT8) 39 ms 14.8 tok/s 2.45 GB Static Buckets (512 / 1024 / 2048) 2.05 GB (Moderate)
iPhone 15 Pro (Apple A17 Pro, 8 GB) Apple MLX (Metal GPU) SmolLM2 1.7B (4-bit Affine) 35 ms 48.2 tok/s 1.32 GB Fully Dynamic ($O(1)$) 3.18 GB (Safe)
iPad Pro M4 (Apple M4, 16 GB) Core ML (ANE Target) Qwen 2.5 3B (INT8) 32 ms 12.2 tok/s 3.95 GB Static Buckets (512 / 1024 / 2048) 12.05 GB (Safe on 16 GB)
iPad Pro M4 (Apple M4, 16 GB) Apple MLX (Metal GPU) Qwen 2.5 3B (4-bit Affine) 26 ms 34.8 tok/s 2.20 GB Fully Dynamic ($O(1)$) 13.80 GB (Safe on 16 GB)

Architectural Best Practices for Mobile AI Runtime Design

Building high-performance, private generative AI software on iOS demands aligning model execution targets with the physical architecture of Apple Silicon:

  1. Partition Workloads by Arithmetic Intensity: Route non-autoregressive encoder models (such as Whisper audio encoders, CLIP image encoders, and embedding models) to the ANE. These models process static input dimensions in single forward passes where the ANE's INT8 matrix throughput excels at minimal battery draw. Route all autoregressive decoder models (LLMs and multimodal text generators) to the GPU.
  2. Standardize on 4-Bit Affine Grouped Quantization: Utilize Apple MLX 4-bit quantization with a group size of 64. This scheme optimizes memory bus utilization during vector-matrix multiplies while maintaining perplexity within 0.15 points of uncompressed models.
  3. Implement Zero-Copy Unified Memory Mapping: Avoid loading model files through intermediate Foundation framework arrays. Use POSIX mmap directly into MTLResourceStorageModeShared buffers to ensure weights are streamed straight from flash storage to unified RAM with zero memory duplication.
  4. Enforce an Absolute 3.5 GB Dirty Memory Budget: To guarantee crash-free background stability on 8 GB iOS hardware, design memory allocation pools so that static weights, intermediate scratchpads, and full-context KV caches never exceed 3.5 GB of resident dirty RAM.

References & Technical Papers

  • Apple Neural Engine: Architecture, Programming, and Performance

    Spencer H. Bryngelson (arXiv:2606.22283, 2026)

  • Orion: Characterizing and Programming Apple's Neural Engine for LLM Training and Inference

    Ramchand Kumaresan (arXiv:2603.06728, 2026)

  • Deploying Transformers on the Apple Neural Engine

    Apple Machine Learning Research (2022)

Local execution with Lapis

Lapis runs these models natively on your iPhone or iPad using Apple MLX and Metal shaders, completely air-gapped with zero remote servers.

App Store