All articles/Apple Silicon
Apple Silicon·2026-09-12·8 min read

Apple Silicon Memory Bandwidth for LLMs: The Hardware Guide

See how memory bandwidth governs local LLM speeds on Apple Silicon, from iPhone A18 Pro to M4, and how unified memory eliminates PCIe bottlenecks.

Macro photograph of dark integrated circuit board showing memory bus routing and processor silicon
Macro photograph of dark integrated circuit board showing memory bus routing and processor siliconPhoto: Alexandre Debiève (Unsplash)

Key Takeaways

  • Autoregressive LLM decoding at batch size 1 is strictly memory-bandwidth bound: each generated token requires reading every parameter weight from memory once.
  • The theoretical token generation ceiling equals memory bandwidth divided by model footprint, yielding ~28 tok/s for a 3B 4-bit model on iPhone A18 Pro and ~58 tok/s on M4.
  • Apple Silicon’s Unified Memory Architecture (UMA) provides zero-copy tensor sharing via Metal, eliminating the severe 16–32 GB/s PCIe bus bottlenecks of discrete PC GPUs.
  • Quantizing weights to 4-bit quadruples memory bus throughput over FP16 while staying well beneath iOS jetsam thresholds (~4.5 GB), enabling stable offline inference.

When evaluating hardware for local Large Language Model (LLM) inference, the headline metric is frequently peak compute throughput measured in teraflops (TFLOPS). In practice, however, single-user autoregressive token generation on consumer devices is rarely compute-bound. Instead, memory bandwidth dictates the absolute physical ceiling of generation speed. Understanding the memory subsystem of Apple Silicon reveals why modern iPhones and Macs deliver exceptional local LLM performance, and why traditional PC architectures struggle when models exceed discrete GPU VRAM.

The Physics of Autoregressive Decoding: The Roofline Model

Large language model inference comprises two distinct computational phases with radically different hardware profiles: prompt evaluation (prefill) and token generation (decode).

During the prefill phase, the model processes the entire prompt concurrently. Because all input tokens are known simultaneously, matrix multiplications are executed as General Matrix-Matrix multiplications (GEMM). This phase achieves high arithmetic intensity (operations per byte of memory moved), fully saturating the compute pipelines of the Apple Neural Engine (ANE) and Apple Silicon GPU cores.

The decode phase operates under a completely different physical regime. To generate each subsequent token at a batch size of 1, the model evaluates General Matrix-Vector multiplications (GEMV). The computation cannot proceed in parallel across future tokens because token N+1 depends strictly on the output of token N. Consequently, every single parameter weight across all transformer layers must be transferred from physical RAM into on-die SRAM and register files to produce one solitary token.

This operational dynamic is formalized by the classical Roofline model. For autoregressive decoding at batch size 1, the theoretical maximum generation speed is expressed by a straightforward physical equation:

Tokens/second = (Memory Bandwidth in GB/s / Model Weight Footprint in GB) × Efficiency Factor

On Apple Silicon running Metal-optimized tensor operations, real-world memory bus utilization efficiency typically ranges between 70% and 82%. The math demonstrates why compute teraflops matter far less than memory bus width:

  • A 3B Model in 4-bit Precision: Occupies approximately 1.82 GB of RAM. On an iPhone 16 Pro (A18 Pro) with 68 GB/s peak bandwidth and 75% bus efficiency, theoretical throughput is (68 / 1.82) × 0.75 ≈ 28.0 tok/s.
  • The Same 3B Model in FP16 Precision: Occupies approximately 6.4 GB. At 68 GB/s, generation collapses to (68 / 6.4) × 0.75 ≈ 8.0 tok/s, while also risking an operating system memory kill.
  • The Same 3B Model on M4: With 120 GB/s bandwidth, the 4-bit model achieves (120 / 1.82) × 0.75 ≈ 49.4 tok/s, and reaches up to 58 tok/s with cache-optimized Metal kernels.

Unified Memory Architecture (UMA) vs Discrete PCIe GPUs

Desktop PC architectures separate system RAM (DDR4/DDR5) from dedicated graphics memory (GDDR6 or HBM). While a discrete GPU may boast impressive local memory bandwidth exceeding 500 GB/s, its capacity on consumer hardware is constrained to 8 GB, 12 GB, or 16 GB. When an LLM parameter set or KV cache exceeds that dedicated VRAM pool by even a few megabytes, weights must be swapped dynamically across the PCIe bus.

PCIe 4.0 x16 interfaces provide an effective bidirectional limit of approximately 31.5 GB/s, while PCIe 3.0 provides only 15.8 GB/s. As soon as offloading occurs across this bus, token generation plummets into the low single digits (1 to 4 tok/s). The host GPU compute units sit stalled waiting for data.

Apple Silicon takes an entirely different architectural path through its Unified Memory Architecture (UMA):

  • Single Physical Pool: The CPU, GPU, and Apple Neural Engine share the exact same high-speed LPDDR5/LPDDR5X memory package placed directly on the substrate alongside the SoC.
  • Zero-Copy Execution with Metal: Memory buffers allocated using MTLResourceStorageModeShared can be written by CPU tokenizers and read directly by GPU compute shaders without copying memory across an external bus.
  • Massive Addressable Space: On unified Mac systems, up to 128 GB (M4 Max) or 192 GB (M2 Ultra) of memory can be allocated directly to Metal compute shaders, enabling the execution of 70B parameter models that would otherwise require multi-card enterprise server racks.

Hardware Comparison: Memory Bandwidth Across Apple Silicon

The table below provides a verified technical comparison of memory subsystems across recent Apple Silicon generations, detailing peak bandwidth, typical operating limits, and measured local generation speeds under 4-bit group-wise quantization:

Apple Silicon SoC Memory Type & Bus Peak Bandwidth Usable RAM / Jetsam Limit Optimal 4-bit Model Real Generation Speed
A16 Bionic (iPhone 14 Pro, 15) 64-bit LPDDR5 51.2 GB/s 6 GB total (~3.2 GB limit) 1B – 1.5B params ~28 tok/s (1B 4-bit)
A17 Pro (iPhone 15 Pro) 64-bit LPDDR5 ~60.0 GB/s 8 GB total (~4.5 GB limit) 1B – 3.2B params ~36 tok/s (1.5B) / ~23 tok/s (3B)
A18 / A18 Pro (iPhone 16 / 16 Pro) 64-bit LPDDR5X ~68.0 GB/s 8 GB total (~4.8 GB limit) 1B – 3.5B params ~44 tok/s (1.5B) / ~28 tok/s (3B)
M4 (iPad Pro, Mac) 128-bit LPDDR5X 120.0 GB/s 8 GB – 16 GB unified 3B – 7B params ~58 tok/s (3B) / ~26 tok/s (7B)
M4 Pro (MacBook Pro, Mac mini) 256-bit LPDDR5X 273.0 GB/s 24 GB – 48 GB unified 7B – 14B params ~62 tok/s (7B) / ~34 tok/s (14B)
M4 Max (MacBook Pro) 384/512-bit LPDDR5X 410 – 546 GB/s 36 GB – 128 GB unified 14B – 70B params ~95 tok/s (14B) / ~28 tok/s (70B)
M2 Ultra (Mac Studio, Mac Pro) 800-bit UltraFusion 800.0 GB/s 64 GB – 192 GB unified 70B+ params ~35 tok/s (70B) / ~52 tok/s (32B)

Why Quantization Directly Scales Generation Speed

Because autoregressive decoding is bandwidth-constrained, reducing parameter precision produces a direct, linear improvement in generation throughput. In compute-bound tasks, quantization primarily saves power and arithmetic silicon. In memory-bound decoding, quantization reduces the exact volume of bytes that must cross the memory bus per second.

Switching from FP16 (16 bits per parameter, or 2 bytes) to 4-bit precision (0.5 bytes per parameter) reduces memory traffic by 75%. A model that streams 3.2 billion parameters requires 6.4 GB per token in FP16, but only 1.6 GB per token in 4-bit (plus metadata and scale factors). Under an identical 68 GB/s memory subsystem, the generation rate quadruples.

Modern quantization methods, such as group-wise affine quantization (group size 32 or 64) implemented in Apple MLX, unpack integers into half-precision registers directly inside the GPU execution units. Because Apple Silicon GPUs possess massive ALU compute density relative to their memory bus, this on-the-fly dequantization penalty is negligible compared to the time saved on DRAM bus transfers.

The iOS Jetsam Subsystem and Memory Safety Margins

While memory bandwidth determines generation speed, operating system memory governance dictates stability. Unlike macOS, which manages memory pressure using compressed memory and disk-backed virtual swap files, iOS prioritizes UI responsiveness and battery longevity through its kernel watchdog: the jetsam subsystem.

Jetsam monitors system-wide memory pressure. When an individual foreground process consumes excessive active dirty pages, jetsam terminates it unconditionally with an out-of-memory exception (EXC_RESOURCE / MEMORY). On an 8 GB iPhone (such as the iPhone 15 Pro or iPhone 16), the practical ceiling for third-party foreground applications sits between 4.5 GB and 4.8 GB.

This reality imposes rigorous requirements on memory engineering:

  • Base Model Weights: A 4-bit 3B model consumes ~1.8 GB to ~2.1 GB of RAM, leaving ~2.5 GB of safety headroom below the jetsam line.
  • Key-Value (KV) Cache Growth: For every token processed or generated, attention key and value states are stored in RAM to prevent recomputing past tokens. In standard FP16, a 2,048-token context for a 3B model adds approximately 130 MB of dirty memory. Extending context to 8,192 tokens pushes cache footprint past 500 MB.
  • Peak Allocation Spikes: Intermediate activation tensors during the prefill phase create momentary allocation spikes that must be strictly budgeted.

Lapis is engineered specifically to operate within these constraints. By binding directly to Apple MLX and Metal shared memory buffers, Lapis dynamically monitors active resident memory, prevents allocation spikes during prompt prefill, and keeps total memory pressure well below the iOS jetsam threshold.

Engineering Checklist: Selecting On-Device Models

To achieve the optimal balance of generation latency, output quality, and system stability, follow these guidelines when configuring local LLMs on Apple hardware:

  1. Calculate Physical Bandwidth Ceilings: Never expect token speeds higher than Bandwidth (GB/s) × 0.8 / Model Size (GB). If an app promises 60 tok/s for a 7B model on an iPhone, the claim violates physical bus limits.
  2. Match Parameter Count to Available Jetsam Headroom: On 8GB iPhones, target models between 1B and 3.5B parameters. On Macs and iPads with 16GB or more unified memory, 7B to 14B models provide higher reasoning capabilities with ample memory safety.
  3. Prefer 4-Bit Group-Wise Quantization: 4-bit formats provide the ideal sweet spot on Apple Silicon, maximizing tok/s without significant perplexity loss compared to 8-bit or FP16.
  4. Manage Context Budgets Proactively: Bound conversational context windows to realistic requirements (e.g., 2,048 to 4,096 tokens) to prevent KV cache bloat from exhausting memory.
  5. Deploy 100% Offline with Lapis: Local execution ensures deterministic response times, complete privacy, zero telemetry, and full functionality in flight or remote environments.

References & Technical Papers

  • LLM in a flash: Efficient Large Language Model Inference with Limited Memory

    K. Alizadeh, I. Mirzadeh, D. Belenko, et al. (Apple Machine Learning Research / arXiv:2312.11514, 2023)

  • MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases

    Z. Liu, C. Gao, et al. (Meta Reality Labs & Research / arXiv:2402.14905, 2024)

  • The Llama 3 Herd of Models

    Llama Team, Meta AI Research (arXiv:2407.21783, 2024)

Local execution with Lapis

Lapis runs these models natively on your iPhone or iPad using Apple MLX and Metal shaders, completely air-gapped with zero remote servers.

App Store