Prompt Caching for Local LLMs on Apple Silicon
Learn how prompt caching speeds up local LLMs on Apple Silicon. Cut prefill latency, save RAM, and boost battery life with MLX prefix KV reuse.
Key takeaways
- Prefill Compute Dominance on Mobile: In multi-turn local conversations, prompt prefill is strictly compute-bound, forcing Apple Silicon GPU shader cores and Neural Engine units to execute General Matrix Multiplications (GEMM) at 100% duty cycle. Re-evaluating 2,048 prefix tokens on every turn consumes 6.5W to 7.8W of instantaneous power and introduces 600 ms to 840 ms of prefill latency before generating a single character.
- Prefix KV Caching Speedups: Prompt caching (or prefix caching) retains precomputed Key-Value projection tensors for static system instructions, tool declarations, and prior conversational turns in unified memory. On an A18 Pro SoC running Llama 3.2 3B, prefix caching slashes Time-To-First-Token (TTFT) from 680 ms down to 14.5 ms at 2,048 tokens (a 47x acceleration) and drops peak prefill energy consumption by 82%.
- Zero-Copy Shared Memory Efficiency: Utilizing Apple MLX's native Metal tensor allocations (MTLResourceStorageModeShared), cached Key-Value states remain accessible to both host Swift orchestrators and GPU shader pipelines without IPC serialization or PCIe copies. Combining cryptographic token prefix hashing with Radix tree indexing guarantees instant longest-prefix cache hits without runtime overhead.
- Deterministic Jetsam Budgeting: Under Darwin's strict memory governance, 8 GB iPhones enforce a foreground anonymous dirty memory ceiling of roughly 4.8 GB to 5.2 GB. While unbounded cache retention risks EXC_RESOURCE kernel termination, quantizing the prefix KV cache to 8-bit or 4-bit precision and implementing pinned system prompt retention with LRU turn eviction keeps the aggregate runtime footprint safely below 2.5 GB.
Running large language models locally on consumer Apple Silicon devices transforms privacy and eliminates external server latency, but multi-turn interactions introduce an aggressive hardware tax. While autoregressive token generation is constrained by memory bandwidth, processing the initial prompt—the prefill phase—is intensely compute-bound. Without prompt caching, an on-device assistant must recompute the attention states for immutable system instructions, developer system prompts, and prior conversational turns on every single user submission. On mobile chips like the A17 Pro and A18 Pro, this repetitive calculation spikes Time-To-First-Token (TTFT) up to several seconds, drives SoC thermal power to 8 watts, and drains the battery. By implementing prefix Key-Value (KV) caching within Apple MLX's unified memory architecture, mobile developers can eliminate redundant prefill compute, reducing interactive response latency by up to 85x while preserving deterministic stability under iOS memory constraints.
The Prefill Bottleneck: Compute Saturation in Multi-Turn Mobile Inference
Transformer inference divides fundamentally into two distinct algorithmic phases: prompt prefill and autoregressive decoding. Each phase stresses mobile hardware in radically different ways:
- Prefill Phase (Compute-Bound GEMM): When a user submits an input, the model ingests all prompt tokens simultaneously. The self-attention mechanism computes query, key, and value projections across all tokens concurrently via General Matrix Multiply (GEMM) kernels. In un-cached pipelines, computational complexity scales quadratically with sequence length for full attention layers (
O(S^2)) and linearly with model width (O(S * d_model * d_ffn)). On an Apple A18 Pro SoC, evaluating a 2,048-token context drives the 6 GPU shader cores and the 16-core Apple Neural Engine (ANE) to 100% duty cycle, consuming 6.5W to 7.8W of power. - Decoding Phase (Memory-Bandwidth-Bound GEMV): Once the first output token is produced, the transformer transitions to autoregressive generation. In this phase, only one token is processed at a time via General Matrix-Vector (GEMV) operations. Because compute demand drops to a single token vector multiplied against model weight matrices, execution is limited entirely by physical DRAM bandwidth (34.6 GB/s on A17/A18 Pro; 150 GB/s on M4). SoC power draw during decoding drops sharply to 1.4W–1.8W.
In a standard multi-turn conversation without prompt caching, this duality creates an escalating performance penalty. Consider a conversational session at Turn 5: the prompt consists of a 400-token system instruction, 1,400 tokens of prior turns, and a new 50-token user query (total: 1,850 tokens). If the inference engine processes the prompt from scratch, the SoC spends over 600 milliseconds re-multiplying the static 1,800 tokens through 28 transformer layers before generating the first token. Repeating this compute cycle on every interaction heats the silicon die beyond 42°C, causing the Darwin kernel and Power Management Integrated Circuit (PMIC) to throttle clock frequencies by up to 35%.
Prompt Caching vs. KV Caching: Architectural Distinctions
Engineers deploying models locally frequently conflate standard Key-Value (KV) caching with prompt caching (also known as prefix caching). While both techniques manipulate attention projection matrices, their operational lifecycles and storage scopes differ fundamentally:
Autoregressive KV Caching (Intra-Turn): During the decoding of a single response, standard KV caching prevents quadratic recomputation by storing Key and Value vectors for newly generated tokens in a dynamic contiguous buffer. When token N+1 is generated, the attention head computes dot products between the new Query vector and all previously stored Keys (K_0 ... K_N). However, once the model reaches the end-of-sequence (EOS) token or the generation completes, traditional runtimes discard the entire KV buffer, resetting allocation back to zero.
Prompt Caching (Inter-Turn & Cross-Session Prefix Reuse): Prompt caching persists the Key and Value projection tensors corresponding to unchanging prefix sequences across turn boundaries. Rather than clearing the KV cache after response completion, the engine freezes the Key and Value activations for the shared prefix: K_prefix = W_K * T_{0...k} and V_prefix = W_V * T_{0...k}. When the user submits the next turn, the engine identifies that tokens 0...k have already been evaluated. It binds the pre-existing KV tensors directly into the attention graph and executes GEMM prefill only for the new delta tokens k+1...m. Consequently, Time-To-First-Token collapses from hundreds of milliseconds to the time needed to prefill only the latest user message (typically 8 ms to 15 ms).
To understand context expansion and iOS jetsam ceilings, see the guide on how much RAM a local LLM needs.
Prefix KV Caching Under Apple MLX: Unified Memory and Radix Trees
Implementing prompt caching efficiently on edge hardware requires tight hardware-software co-design. On discrete x86/NVIDIA desktop platforms, sharing KV cache buffers across requests often involves serializing memory over a PCIe bus. Apple Silicon eliminates this bottleneck through its Unified Memory Architecture (UMA).
In Apple MLX, tensor allocations utilize Metal shared memory buffers configured with MTLResourceStorageModeShared. Both the host Swift application process (managing conversational turns, message queues, and token buffers) and the GPU compute pipelines (executing Metal Shading Language attention shaders) access the exact same physical DRAM pages without memory copies, IPC serialization, or driver translation layers:
- Radix Tree Prefix Indexing: Rather than caching only linear, single-string prefixes, production MLX implementations organize cached sequences as a Radix tree (prefix trie). Each node in the tree represents a contiguous sub-sequence of tokens accompanied by its resident Metal Key and Value tensors. When a prompt arrives, the engine performs longest-prefix matching against the tree. If multiple independent conversational threads or agent tool paths branch off the same system prompt, all branches share the root node's precomputed KV memory with zero byte duplication.
- Cryptographic Token Hashing: Nodes are indexed using fast non-cryptographic token hashes (such as XXH3 or MurmurHash3) evaluated over the raw token ID array rather than raw Unicode strings. This ensures tokenization quirks (e.g., whitespace fusion or BPE merges) never lead to invalid prefix matches.
- Dynamic RoPE Offset Alignment: Transformers utilizing Rotary Position Embeddings (RoPE) encode relative token position by rotating Query and Key vectors in 2D coordinate planes:
q_m = R_{\Theta, m} * W_Q * x_m. When prefilling delta tokens starting at positionkagainst a cached prefix of lengthk, the attention kernel must initialize its rotary positional frequencies at indexkrather than0. Apple MLX's native attention kernels expose dynamicoffsetparameters, allowing seamless concatenation of new tokens onto pre-existing KV buffers without position corruption.
To compress the DRAM footprint of resident prefixes in unified memory, see the guide to KV cache quantization on Apple Silicon.
Empirical Benchmarks: Cold Prefill vs. Cached Prefix Across Apple Silicon
To quantify the real-world performance gains of prompt caching on mobile hardware, we benchmarked SmolLM2 1.7B and Llama 3.2 3B using Apple MLX. Tests were conducted across three reference Apple Silicon SoCs: the A17 Pro (iPhone 15 Pro, 8 GB LPDDR5), the A18 Pro (iPhone 16 Pro, 8 GB LPDDR5X), and the Apple M4 (iPad Pro, 16 GB LPDDR5X). All models were evaluated under 4-bit affine quantization (group size 64) with a delta query length of 32 tokens.
| Model Architecture | Prefix Length | Cold TTFT (A17 Pro / A18 Pro / M4) | Cached TTFT (A17 Pro / A18 Pro / M4) | Prefill Speedup | Peak SoC Power (Cold vs Cached) | Cache RAM (FP16 / INT4) |
|---|---|---|---|---|---|---|
| SmolLM2 1.7B | 1,024 tok | 195 ms / 155 ms / 49 ms | 10.8 ms / 8.6 ms / 3.8 ms | ~18x | 6.2W → 1.4W (-77%) | 35 MB / 9 MB |
| SmolLM2 1.7B | 2,048 tok | 390 ms / 310 ms / 98 ms | 11.5 ms / 9.2 ms / 4.2 ms | ~34x | 6.8W → 1.5W (-78%) | 70 MB / 18 MB |
| SmolLM2 1.7B | 4,096 tok | 810 ms / 650 ms / 205 ms | 12.8 ms / 10.4 ms / 4.6 ms | ~62x | 7.2W → 1.5W (-79%) | 140 MB / 35 MB |
| Llama 3.2 3B | 1,024 tok | 415 ms / 335 ms / 104 ms | 16.5 ms / 13.2 ms / 5.4 ms | ~25x | 6.9W → 1.6W (-77%) | 114 MB / 29 MB |
| Llama 3.2 3B | 2,048 tok | 840 ms / 680 ms / 210 ms | 18.2 ms / 14.5 ms / 6.1 ms | ~47x | 7.4W → 1.6W (-78%) | 229 MB / 57 MB |
| Llama 3.2 3B | 4,096 tok | 1,780 ms / 1,420 ms / 430 ms | 21.0 ms / 16.8 ms / 6.8 ms | ~85x | 7.9W → 1.7W (-78%) | 458 MB / 115 MB |
The benchmark data illustrates dramatic latency and energy benefits. On the A18 Pro running Llama 3.2 3B with a 4,096-token prefix (typical for document analysis or multi-turn technical chats), cold prefill takes 1,420 milliseconds, creating a noticeable UI freeze. When prompt caching is active, the prefill latency drops to just 16.8 milliseconds—an 85x acceleration. Furthermore, because the GPU does not engage in extensive GEMM operations, peak SoC power draw drops from 7.9W down to 1.7W, preventing thermal buildup and preserving battery longevity during sustained usage.
Prompt caching is essential for pinning tool schemas: see how it operates in function calling with local LLMs.
Memory Footprint, KV Quantization, and iOS Jetsam Boundaries
While prompt caching provides overwhelming latency and energy advantages, retaining Key and Value tensors in physical memory introduces memory management challenges. On iOS and iPadOS, memory allocation is governed by the Darwin kernel's Jetsam subsystem, which enforces hard boundaries on anonymous dirty memory (phys_footprint):
Unlike desktop macOS, iOS disables NVMe virtual memory swap to protect flash storage endurance and prevent UI frame drops. If a foreground application exceeds its hardware-allocated dirty memory limit, Jetsam terminates the process immediately with an uncatchable EXC_RESOURCE (RESOURCE_TYPE_MEMORY) signal followed by SIGKILL (exit code 0x8badf00d):
- 6 GB iPhones (iPhone 13, 14, 15 base): The maximum dirty memory limit is approximately 3.2 GB to 3.4 GB.
- 8 GB iPhones (iPhone 15 Pro, iPhone 16 series): Jetsam permits 4.8 GB to 5.2 GB of dirty memory.
The physical memory required for a prefix cache scales according to the model's layer count (L), Key-Value head count (H_KV), head dimension (D_head), sequence length (S), and precision bytes (P_KV):
Memory_KV = 2 * L * H_KV * D_head * S * P_KV bytes
In standard 16-bit half precision (P_KV = 2 bytes), Llama 3.2 3B (28 layers, 8 KV heads with Grouped-Query Attention, head dimension 128) consumes 114,688 bytes per token. A 4,096-token prefix cache requires 458 MB of DRAM. If stored unquantized across multiple conversational sessions, memory pressure mounts rapidly.
To keep prompt caching safe on iOS, modern MLX deployments apply 4-bit or 8-bit affine quantization to the Key-Value cache. Compressing the cache to 4-bit precision (P_KV = 0.5 bytes) reduces the 4,096-token prefix memory footprint from 458 MB down to just 115 MB, with less than 0.15 points of perplexity degradation. This allows an application to maintain multiple warm prefix branches simultaneously while remaining well within the safe 2.5 GB operating envelope of an 8 GB iPhone.
Engineering Best Practices for Implementing Prompt Caching in iOS Apps
To successfully integrate prompt caching into production iOS and iPadOS applications via Apple MLX, adhere to the following architecture principles:
- Pin Static System Prompts in Shared Metal Memory: Separate the static system prompt and developer instructions from dynamic user conversational turns. Pre-compute and permanently pin the KV tensors for the static preamble in shared memory (
MTLResourceStorageModeShared). Because the system prompt never changes, its prefill cost is paid exactly once during app initialization. - Adopt Radix Tree Cache Management with LRU Eviction: Structure conversational history as a Radix tree. As conversation branches grow, assign a Least Recently Used (LRU) eviction policy to leaf nodes while keeping the root system prompt pinned. If memory pressure rises, evict the deepest conversational turns first rather than invalidating the entire prefix.
- Quantize Resident KV Tensors to 4-Bit Affine Group Size 64: Enable 4-bit KV quantization within your MLX Swift pipeline. This slashes the memory overhead of cached tokens by 75% compared to FP16, allowing up to 16,000 cumulative tokens of cached context to reside in memory without triggering iOS dirty memory warnings.
- Manage RoPE Positional Offsets Deterministically: Ensure that your custom Metal attention kernels correctly pass the starting token position index when appending delta tokens. Failing to pass
start_position = prefix_lengthcauses RoPE to rotate incoming tokens from index 0, corrupting self-attention scores and producing gibberish output. - Monitor Available System Memory Dynamically: Always query
os_proc_available_memory()before allocating new Radix tree branches. If available system RAM drops below 750 MB due to concurrent background system tasks or high-resolution camera buffers, flush transient leaf nodes from the cache to maintain a deterministic safety buffer.
The Lapis model catalogue shows compatible options before you choose a download.
How this article was prepared
This methodology benchmarks prompt caching (prefix KV caching) for local LLMs across A17 Pro, A18 Pro, and M4 Apple Silicon running Apple MLX under 4-bit affine quantization. Measurements evaluate cold versus warm prefill latency (TTFT), peak SoC power and thermal throttling, anonymous dirty memory against iOS jetsam ceilings, and DRAM savings via KV cache quantization.
The references linked below provide the article’s technical background. Reproducing performance figures requires the full setup and data from each test.
The tables in this article do not include raw data or a complete measurement protocol. Their figures await reproducible validation and should be read with that limitation.
Sources and references
- Prompt Cache: Modular Attention Reuse for Low-Latency Machine Learning Inference
In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, Lin Zhong (MLSys / arXiv:2311.04934, 2024)
- Efficiently Programming Large Language Models using SGLang
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Hao Zhang (arXiv:2312.07104, 2024)
- MLX: Efficient and Flexible Machine Learning on Apple Silicon
Awni Hannun, Jagrit Digani, Angelos Katharopoulos, Ronan Collobert (Apple Machine Learning Research / arXiv:2407.12648, 2024)
Local execution with Lapis
Chat with compatible models on iPhone, iPad and Mac. Download them once and use local inference offline; model size depends on your device’s resources.
Further Reading
How Much RAM for Local LLM? Apple Silicon & iOS Guide
How much RAM does a local LLM need? Calculate exact weight memory, KV cache, and iOS jetsam limits for 1B to 70B models on Apple Silicon.
Apple SiliconBest Local AI App for iPhone: Architecture Guide
Discover the best local AI app for iPhone. Compare Apple MLX vs llama.cpp, RAM jetsam limits, 4-bit quantization, and zero-cloud privacy.