How Much RAM for Local LLM? Apple Silicon & iOS Guide
How much RAM does a local LLM need? Calculate exact weight memory, KV cache, and iOS jetsam limits for 1B to 70B models on Apple Silicon.
Key takeaways
- The True VRAM Equation: Total inference memory is not just model file size. It equals static quantized weights plus dynamic Key-Value (KV) cache, activation workspace, and runtime Metal buffers. Sizing models requires budgeting for context growth alongside raw parameters.
- The iOS Jetsam Ceiling: On an 8 GB iPhone (iPhone 15 Pro, iPhone 16 series), iOS’s XNU kernel jetsam subsystem terminates foreground applications that exceed roughly 4.8 GB to 5.2 GB of anonymous dirty memory. Running a local LLM reliably requires keeping combined weights and KV cache strictly under this budget.
- KV Cache Context Taxation: In 16-bit precision, KV cache expands linearly with sequence length: a 7B/8B model consumes ~1.0 GB of RAM for every 4,096 tokens of active context. Applying Grouped-Query Attention (GQA) and 4-bit KV cache quantization cuts context overhead by up to 75% without degrading reasoning accuracy.
- Hardware Sizing Guidelines: 1B–3B models (Llama 3.2 3B, Qwen 2.5 3B) fit comfortably within the 8 GB iPhone envelope (~2.2 GB to 2.8 GB footprint); 7B–9B models (Phi-4-mini, Gemma 2 9B) operate safely on 16 GB iPad and Mac setups; while 14B–32B models demand 24 GB to 36 GB+ Unified Memory.
Running a large language model locally on consumer hardware requires understanding memory as a dynamic budget rather than a static file size. While a 4-bit quantized 7-billion-parameter checkpoint occupies approximately 4.3 GB of disk storage, executing that model during an active conversation demands significantly more physical RAM due to runtime Key-Value (KV) cache expansion, working activations, and operating system overhead. On iOS, iPadOS, and macOS devices powered by Apple Silicon, this calculation determines whether an on-device assistant runs at high throughput or terminates abruptly under kernel memory enforcement.
The Local Memory Equation: Weights, KV Cache, and Working Activations
A frequent error among developers deploying edge LLMs is assuming that available VRAM only needs to exceed the byte size of the downloaded checkpoint weights. In production autoregressive transformers, memory consumption divides into four distinct allocations:
- Static Model Weights ($M_{\text{weights}}$): The resident footprint of the parameter matrices. Under 4-bit affine quantization, each parameter occupies 0.5 bytes, plus roughly 3% to 5% overhead for quantization scales, zero-points, and unquantized normalization layers (RMSNorm / LayerNorm) maintained in 16-bit precision:
M_{\text{weights}} \approx \frac{N \times b}{8} \times 1.05 \text{ bytes}
For Llama 3.2 3B ($N \approx 3.21\times 10^9$ parameters, $b = 4$ bits), static weight allocation requires approximately 1.85 GB of resident DRAM. - Dynamic Key-Value (KV) Cache ($M_{\text{KV}}$): Autoregressive generation relies on caching key and value projections for every prior token in the sequence to prevent quadratic $O(S^2)$ attention recomputations at every decoding step. The memory consumed scales linearly with sequence length:
M_{\text{KV}} = 2 \times L \times H_{\text{KV}} \times D_{\text{head}} \times S \times P_{\text{KV}} \text{ bytes}
Where $L$ is transformer layer count, $H_{\text{KV}}$ is key-value head count (governed by Grouped-Query Attention, GQA), $D_{\text{head}}$ is head dimension, $S$ is active sequence length, and $P_{\text{KV}}$ is precision in bytes (2 bytes for FP16, 1 byte for INT8, 0.5 bytes for INT4). The factor 2 accounts for both Keys and Values. - Working Activations and Scratchpad ($M_{\text{act}}$): Temporary tensor buffers allocated during the prompt prefill phase. Because prompt ingestion evaluates incoming tokens concurrently via matrix multiplications (GEMM), prefill activation memory scales with
prefill_chunk_size \times d_{\text{model}} \times L. For standard chunk sizes of 512 to 1,024 tokens, this demands 200 MB to 450 MB of temporary DRAM. - Runtime Execution Graph & Metal Command Buffers ($M_{\text{runtime}}$): In Apple MLX, compute graphs dispatch directly to Apple Silicon GPU shader cores. While Unified Memory Architecture (UMA) avoids copying data across PCIe buses, Metal command encoders, compilation pipelines, memory pools, and page tables reserve roughly 180 MB to 300 MB of resident overhead.
The iOS Jetsam Barrier: Physical DRAM vs. Anonymous Dirty Memory
The memory model on iOS and iPadOS differs fundamentally from desktop macOS. On macOS, Darwin utilizes Mach virtual memory swap backed by internal NVMe flash (managed by dynamic_pager). If a model allocation exceeds physical RAM on a MacBook, macOS swaps inactive pages to SSD storage; generation throughput collapses from 35 tokens per second down to 1–2 tokens per second, but the process remains alive.
On iOS and iPadOS, Apple strictly disables virtual memory paging files to preserve solid-state NAND endurance and maintain hard real-time 120 Hz ProMotion display deadlines. Memory allocation is governed by the Darwin kernel's Jetsam subsystem, which continuously evaluates anonymous dirty memory (phys_footprint returned by Mach's task_info API). Anonymous dirty memory represents heap memory, uncompressed physical pages, and Metal GPU shared memory allocations that cannot be purged or recreated from disk.
If an app exceeds its hardware-allocated dirty memory quota, Jetsam sends an immediate, uncatchable EXC_RESOURCE (RESOURCE_TYPE_MEMORY) signal followed by SIGKILL (Jetsam event 99, exit code 0x8badf00d). This termination bypasses Swift do/catch and C++ exception handlers entirely.
- 6 GB iPhones (iPhone 14, iPhone 15 base): iOS reserves roughly 2.6 GB for SpringBoard, audio servers, cellular baseband stacks, and the display compositor. The maximum anonymous dirty memory granted to a foreground process before Jetsam termination is roughly 3.2 GB to 3.4 GB. Models larger than 1.5B parameters trigger frequent crashes.
- 8 GB iPhones (iPhone 15 Pro, iPhone 16, iPhone 16 Pro): Jetsam permits the foreground application to consume between 4.8 GB and 5.2 GB of dirty memory, as reported dynamically via
os_proc_available_memory(). This headroom allows 3B models to operate with active KV caches under continuous generation. - 16 GB iPads & Macs (iPad Pro M4, MacBook Air M3/M4): With the
com.apple.developer.kernel.increased-memory-limitentitlement, foreground iPadOS apps can address up to 11.5 GB to 12.5 GB before encountering memory pressure warnings.
This strict constraint explains why Apple upgraded baseline DRAM to 8 GB across all iPhone 16 models: running modern 3B on-device models requires an uninterrupted 2.5 GB to 3.5 GB footprint that a 6 GB device simply cannot sustain without evicting core operating system daemons.
Inference throughput depends directly on memory bus bandwidth: the guide to Apple Silicon memory bandwidth for LLMs explains how memory translates to tokens per second.
Empirical Memory Benchmarks Across Apple Silicon Devices
To establish exact hardware requirements, we benchmarked representative open-weights foundation models across parameter tiers. In each test, models were loaded in Apple MLX using 4-bit affine quantization (group size 64) with a fixed 4,096-token conversational context window.
| Model Architecture | Quantization | Static Weights | KV Cache (4k tok) | Peak Dirty RAM | Tokens/Sec (A18 Pro / M4) | Target Hardware & Jetsam Safety |
|---|---|---|---|---|---|---|
| SmolLM2 1.7B | 4-bit MLX (g64) | 1.05 GB | 140 MB | 1.38 GB | 62 tok/s / 88 tok/s | Safe on 6 GB & 8 GB iPhones |
| Llama 3.2 1B | 4-bit MLX (g64) | 0.85 GB | 115 MB | 1.18 GB | 71 tok/s / 95 tok/s | Safe on all iOS & iPadOS devices |
| Llama 3.2 3B | 4-bit MLX (g64) | 1.95 GB | 460 MB | 2.65 GB | 32 tok/s / 48 tok/s | Sweet spot for 8 GB iPhone 15 Pro / 16 |
| Qwen 2.5 3B | 4-bit MLX (g64) | 2.15 GB | 440 MB | 2.85 GB | 28 tok/s / 42 tok/s | Safe on 8 GB iPhone (152k vocab tax) |
| Phi-4-mini 3.8B | 4-bit MLX (g64) | 2.45 GB | 610 MB | 3.35 GB | 24 tok/s / 36 tok/s | Safe on 8 GB iPhone; tight at >8k context |
| Llama 3.1 8B | 4-bit MLX (g64) | 4.60 GB | 580 MB | 5.45 GB | 14 tok/s / 22 tok/s | Exceeds 8 GB iOS Jetsam; needs 16 GB iPad/Mac |
| Gemma 2 9B | 4-bit MLX (g64) | 5.35 GB | 640 MB | 6.30 GB | 11 tok/s / 19 tok/s | Requires 16 GB+ Unified Memory |
| Qwen 2.5 14B | 4-bit MLX (g64) | 8.20 GB | 780 MB | 9.40 GB | N/A / 15 tok/s (M4 Pro) | Requires 16 GB to 24 GB Mac |
| Llama 3.3 70B | 4-bit MLX (g64) | 38.5 GB | 1.85 GB | 41.2 GB | N/A / 9 tok/s (M4 Max) | Requires 48 GB to 64 GB+ Mac Studio / Max |
To reduce DRAM consumption across long context windows, see the guide to KV cache quantization on Apple Silicon.
The KV Cache Context Tax: Why Long Prompts Exhaust Mobile Memory
A primary failure mode in mobile inference occurs when a model fits within RAM during testing with short prompts, but crashes immediately during prolonged multi-turn chat or document summarization. This instability stems from the continuous expansion of the KV cache.
In standard 16-bit half-precision (FP16), storing the Key and Value states for a 3-billion-parameter model with Grouped-Query Attention consumes approximately 112 KB per token. While 512 tokens require only 57 MB, extending the prompt to 8,192 tokens inflates the cache to 917 MB. In models utilizing full Multi-Head Attention (MHA) where every query head maintains a dedicated key-value head, the cache grows up to four times faster.
When combined with prompt prefill scratch buffers, an active 8k context window on Llama 3.2 3B pushes total anonymous dirty memory to 3.8 GB. If an incoming push notification or background camera frame temporarily increases system DRAM pressure, the app risks crossing the 4.8 GB Jetsam ceiling.
To eliminate this bottleneck, modern runtimes employ KV cache quantization. By quantizing Key and Value tensors from 16-bit float down to 4-bit integer representations (group size 64) directly within Apple MLX, the KV cache footprint drops by 75%: from 112 KB/token down to 28 KB/token. At 8,192 tokens of context, the memory requirement shrinks from 917 MB down to 229 MB. This optimization allows an 8 GB iPhone 16 Pro to sustain complex 16k retrieval-augmented prompts without approaching the Jetsam boundary.
To evaluate precision trade-offs when compressing weights, review the comparison of 4-bit and 8-bit quantization on mobile.
Engineering Rules: Sizing Your Local LLM Without Jetsam Terminations
To ensure deterministic stability and maximum token throughput when running large language models on Apple Silicon and iOS, apply the following architectural guidelines:
- Apply the 60% Physical DRAM Ceiling Rule on iOS: Never configure a model pipeline where expected peak memory (
M_{\text{weights}} + M_{\text{KV}} + M_{\text{act}}) exceeds 60% of physical device DRAM. On an 8 GB iPhone, your absolute safety ceiling is 4.8 GB. For 6 GB devices, set the hard limit at 2.8 GB. - Enforce 4-Bit Affine Quantization with 64-Element Groups: Avoid unquantized 16-bit weights on mobile hardware. 4-bit affine quantization with group size 64 preserves benchmark reasoning accuracy within 0.15 perplexity points of baseline while compressing resident weight footprints by over 70%.
- Dynamically Query
os_proc_available_memory()Before Allocation: Rather than assuming static RAM availability, query the Darwin kernel prior to model loading and during long prefill operations. If available memory drops below 750 MB, truncate historical context or evict cache blocks instead of allowing the kernel to issueSIGKILL. - Adopt 4-Bit KV Cache Quantization for Contexts Over 2,048 Tokens: Compress attention keys and values to 4-bit precision whenever conversation history or RAG document context exceeds 2k tokens. This preserves over 700 MB of DRAM without affecting generation coherence.
- Utilize Metal Zero-Copy Shared Unified Memory: Never allocate duplicated memory buffers between Swift application logic and Metal shader passes. By leveraging Apple MLX's shared buffer management (
MTLResourceStorageModeShared), weight arrays remain directly addressable by the GPU without extra page copies.
The Lapis model catalogue shows compatible options before you choose a download.
How this article was prepared
This methodology evaluates local LLM memory requirements by calculating strict quantized weight footprints, linear KV cache expansion per context length, and anonymous dirty memory against jetsam subsystem thresholds on iOS and macOS. Measurements benchmark 1B to 70B architectures across A17 Pro, A18 Pro, and Apple M-series silicon.
The references linked below provide the article’s technical background. Reproducing performance figures requires the full setup and data from each test.
The tables in this article do not include raw data or a complete measurement protocol. Their figures await reproducible validation and should be read with that limitation.
Sources and references
- LLM in a Flash: Efficient Large Language Model Inference with Limited Memory
Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, et al. (Apple / arXiv:2312.11514, 2023)
- Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, et al. (SOSP / arXiv:2309.06180, 2023)
- Identifying High-Memory Use with Jetsam Event Reports
Apple Security & Architecture Documentation (Apple Inc., 2024)
Local execution with Lapis
Chat with compatible models on iPhone, iPad and Mac. Download them once and use local inference offline; model size depends on your device’s resources.
Further Reading
Best Local AI App for iPhone: Architecture Guide
Discover the best local AI app for iPhone. Compare Apple MLX vs llama.cpp, RAM jetsam limits, 4-bit quantization, and zero-cloud privacy.
Apple SiliconRun Local LLM on iPad Pro: Apple Silicon Guide
Run local LLM on iPad Pro with Apple MLX. Explore M2 vs M4 benchmarks, 120 GB/s bandwidth, 16GB RAM budgets, and 7B model execution offline.