All articles/Apple Silicon
Apple Silicon·2026-03-04·6 min read

How to Run LLMs Locally on iPhone (No Cloud, 100% Offline)

Learn how to run 1.5B to 7B LLMs directly on your iPhone with Apple Silicon, MLX 4-bit quantization, and Metal shaders without internet.

Apple Silicon iPhone and MacBook workstation with minimalist dark desk layout
Apple Silicon iPhone and MacBook workstation with minimalist dark desk layoutPhoto: Tyler Lastovich (Unsplash)

Key Takeaways

  • Modern iPhones with 8GB RAM (A17 Pro, A18, A18 Pro) can run 4-bit quantized models up to 3B parameters entirely in memory.
  • Apple’s Unified Memory Architecture allows the CPU and GPU to read model weights from the same memory pool with zero PCIe transfer bottleneck.
  • Group-wise 4-bit quantization reduces model footprint by ~70% while preserving over 96% of original reasoning capabilities.
  • Lapis delivers turnkey local execution with direct Hugging Face imports, Metal compute shaders, and full airplane mode capability.

Running a high-performing language model on a battery-powered device with 8GB of unified memory was considered impractical until two breakthroughs converged: low-bit quantization frameworks tailored to Apple Silicon, and efficient 1B-to-3B parameter distilled models.

The Architectural Shift: Why Apple Silicon Changes Everything

Traditional computer architectures separate system RAM from dedicated GPU VRAM. Whenever a matrix multiplication takes place, tensor weights must travel over a PCIe bus. On mobile, this creates thermal throttles and extreme battery drain.

Apple Silicon solves this at the silicon die level through Unified Memory Architecture (UMA). The CPU, GPU cores, and Neural Engine share a single high-bandwidth memory fabric. On an A17 Pro or A18 Pro chip, memory bandwidth exceeds 60 GB/s. When an LLM generates tokens, the GPU reads weights directly in place without intermediate copying.

The iOS Memory Budget: Understanding Jetsam Limits

While an iPhone 15 Pro, iPhone 16, or iPhone 16 Pro packs 8GB of physical LPDDR5X RAM, third-party apps cannot allocate all 8GB. The iOS kernel enforces aggressive limits via the jetsam subsystem:

  • System Reserve: iOS reserves roughly 2.5 GB to 3.0 GB for SpringBoard, audio services, cellular daemons, and system caching.
  • Foreground App Limit: A foreground application can safely hold between 3.8 GB and 4.8 GB before receiving low-memory warnings or termination signals.
  • The Sweet Spot: Quantized 4-bit models between 1.0B and 3.5B parameters require between 900 MB and 2.4 GB of RAM, leaving generous headroom for KV cache, UI rendering, and iOS background services.

4-Bit Quantization: Compressing Weights Without Losing Wit

In standard FP16 (16-bit floating point), every model parameter consumes 2 bytes. A 3-billion parameter model requires 6 GB of RAM just to sit idle in memory, exceeding the safe iOS threshold. Through 4-bit quantization (such as AWQ or MLX group-wise quantization), each weight is compressed to 0.5 bytes:

Model Architecture Parameter Count Format RAM Footprint Tokens / Sec (A18 Pro)
Qwen 2.5 0.5B 490M 4-bit MLX ~450 MB ~68 tok/s
DeepSeek-R1 Distill 1.5B 1.54B 4-bit MLX ~1.15 GB ~34 tok/s
Qwen 2.5 3B 3.09B 4-bit MLX ~2.10 GB ~22 tok/s
SmolVLM2 2.2B (Vision) 2.20B 4-bit MLX ~1.70 GB ~18 tok/s

Step-by-Step: Running Models with Zero Configuration

In the past, running local AI on iOS required compiling C++ binaries like llama.cpp via Xcode, installing TestFlight profiles, and sideloading raw GGUF files. Lapis removes all friction:

  1. Download Lapis from the App Store: No developer account or sideloading certificates needed.
  2. Choose Your Model: Pick from verified local models (Qwen 2.5 for general coding and prose, DeepSeek-R1 for chain-of-thought logic, SmolVLM2 for local image OCR) or import any compatible MLX repository directly from Hugging Face.
  3. Switch to Airplane Mode: Turn off Wi-Fi and Cellular. Prompt the model. Generation starts immediately with sub-millisecond first-token latency and zero network activity.

Battery and Thermals in Daily Practice

A common misconception is that local inference rapidly depletes the iPhone battery. Because Apple Silicon switches off GPU compute cores immediately after generation concludes, idle power draw is strictly 0% CPU. Generating a 300-word response takes roughly 8 to 12 seconds, consuming less than 0.3% of total battery charge on an iPhone 16 Pro.

References & Technical Papers

  • MLX: An efficient machine learning framework for Apple silicon

    A. Hannun, J. Jagtap, et al. (Apple Machine Learning Research, 2023)

  • AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration

    J. Lin, J. Tang, et al. (MLSys 2024 / arXiv:2306.00978)

  • Apple Silicon Unified Memory Architecture Developer Guide

    Apple Developer Documentation (2024)

Local execution with Lapis

Lapis runs these models natively on your iPhone or iPad using Apple MLX and Metal shaders, completely air-gapped with zero remote servers.

App Store