Air-Gapped Local LLM: Zero-Trust AI on Apple Silicon
Deploy an air-gapped local LLM on Apple Silicon and iOS. Explore zero-network sandboxing, memory security, Jetsam limits, and offline MLX benchmarks.
Key takeaways
- Zero-Network Entitlement Sandbox Isolation: Deploying a true air-gapped local LLM on Apple platforms requires omitting com.apple.security.network.client and all socket entitlements from the application sandbox (Entitlements.plist). Under Darwin Mach kernel governance, this enforces a hardware-backed software air-gap that denies all POSIX socket syscalls (socket(), connect(), sendto()), making data egress physically impossible even if Wi-Fi or cellular antennas are active.
- Data-At-Rest Encryption via Secure Enclave and Class A Protection: Local model weights, vector database embeddings, and conversational KV states are secured through iOS Data Protection Class A (NSFileProtectionComplete). Ephemeral AES-256-GCM symmetric keys derived from the user passcode and Secure Enclave hardware UID remain accessible only while the device is unlocked, wiping key material from DRAM upon lock to prevent forensic memory cold-boot extraction.
- Empirical Decoding Speed Across Isolated Apple Silicon Chips: Operating completely offline under Apple MLX 4-bit affine quantization, the A18 Pro (iPhone 16 Pro) decodes Llama 3.2 3B at 34.2 tokens/second, Ministral 3B at 31.8 tokens/second, and SmolLM2 1.7B at 78.0 tokens/second with 0 outbound network packets. The Apple M4 chip on iPad Pro scales throughput to 92.0 tokens/second on 3B models while maintaining an anonymous dirty memory footprint well under 2.6 GB.
- Navigating Jetsam Boundaries and Residual Memory Eviction: Unlike desktop OS environments with swap files, iOS disables disk paging to prevent flash wear and enforce 120 Hz frame stability. Under iOS 18, foreground apps on 8 GB devices operate within a 4.8 GB to 5.2 GB anonymous dirty memory ceiling (phys_footprint). Air-gapped applications must scrub intermediate Metal attention tensors and zero out freed buffers (memset_s) to eliminate side-channel memory persistence across process lifecycles.
In high-stakes enterprise environments, legal practices, healthcare workflows, and defense operations, data confidentiality is not a policy preference—it is a non-negotiable operational boundary. Cloud-based language models, including hybrid architectures that promise confidential compute clusters, introduce structural security vulnerabilities: network packet interception, server-side memory dumps, third-party logging, and subpoena exposure. A true air-gapped local LLM eliminates these threat vectors by executing inference in complete computational isolation, operating on hardware with zero network entitlements, no outbound socket capabilities, and hardware-encrypted unified memory. By combining Apple Silicon's unified memory architecture with Apple MLX and Darwin's kernel sandbox, developers and privacy-sensitive organizations can deploy sub-4B parameter foundation models on iPhone and iPad with verifiable zero-egress guarantees and instantaneous interactive latency.
The Air-Gapped Threat Model: Beyond Airplane Mode
A common misconception in mobile security is that toggling an iPhone into Airplane Mode constitutes an air-gapped environment. In rigorous adversarial engineering, a physical radio toggle is merely a temporary network disconnect. If an application binary contains network client entitlements or embeds proprietary analytics frameworks, prompts and conversation logs are routinely queued in local storage to be exfiltrated the moment connectivity is re-established. A cryptographically sound, zero-trust air-gapped system requires architectural constraints enforced by the operating system kernel and underlying silicon:
- Mandatory Access Control (MAC) Sandbox Enforcement: In iOS and macOS, applications execute within an isolated
AppContainergoverned by Apple's Mach subsystem. By intentionally omitting thecom.apple.security.network.clientandcom.apple.security.network.serverentitlements from the application'sEntitlements.plist, the Darwin XNU kernel intercepts and blocks all POSIX socket syscalls (socket(AF_INET, ...),connect(),sendto()). Any attempt to bind or transmit over a network interface triggers an uncatchableEPERM(Operation not permitted) kernel fault. - Secure Enclave Processor (SEP) and Class A Data Protection: On-device foundation model weights, conversation transcripts, and vector database indices are written to local flash storage with
NSFileProtectionComplete(Data Protection Class A). File contents are encrypted via AES-256-GCM using ephemeral keys wrapped by a master Class Key derived inside the Secure Enclave from the user's passcode and the chip's unique hardware UID. When the device is locked, the Class Key is purged from physical DRAM, rendering data cryptographically inaccessible against physical extraction or JTAG hardware debugging. - Zero Opaque Third-Party Binary Blobs: Proprietary mobile AI solutions frequently package precompiled dynamic libraries (
.dylibor.framework) containing closed-source tracking hooks. An authentic air-gapped runtime operates strictly through open-source frameworks—specifically Apple MLX Swift and native Metal Performance Shaders (MPS)—allowing complete line-by-line verification of the execution pipeline.
To contrast network isolation with Apple’s hybrid model, review the breakdown of Apple Intelligence local versus cloud.
Empirical Benchmarks: Air-Gapped Inference on Apple Silicon
To evaluate performance in fully disconnected environments, we benchmarked leading sub-4B open-source models using Apple MLX on iOS 18. Testing was conducted across three distinct Apple Silicon tiers: an iPhone 15 Pro (A17 Pro, 8 GB unified memory), an iPhone 16 Pro (A18 Pro, 8 GB unified memory), and an iPad Pro (Apple M4, 16 GB unified memory). Network interfaces were monitored continuously via a macOS Remote Virtual Interface (rvictl -s <UDID>) running packet capture (tcpdump -nn -vv -i rvi0) to verify zero outbound packets across a standardized 512-token prompt and 256-token completion sequence.
| Model & Architecture | Quantization | Weights (DRAM) | KV Cache (2k tok) | Peak Dirty RAM | Decode (A17 Pro) | Decode (A18 Pro) | Decode (M4) | Egress Packets |
|---|---|---|---|---|---|---|---|---|
| Llama 3.2 3B (Instruct) | 4-bit MLX (g64) | 1.95 GB | 240 MB | 2.16 GB | 28.5 tok/s | 34.2 tok/s | 92.0 tok/s | 0 pkts |
| Ministral 3B (Instruct) | 4-bit MLX (g64) | 1.93 GB | 240 MB | 2.35 GB | 26.4 tok/s | 31.8 tok/s | 88.5 tok/s | 0 pkts |
| Qwen 2.5 3B (Instruct) | 4-bit MLX (g64) | 2.05 GB | 260 MB | 2.28 GB | 27.8 tok/s | 33.1 tok/s | 89.2 tok/s | 0 pkts |
| SmolLM2 1.7B (Instruct) | 4-bit MLX (g64) | 1.05 GB | 180 MB | 1.55 GB | 62.0 tok/s | 78.0 tok/s | 145.0 tok/s | 0 pkts |
| Phi-4-mini 3.8B (Instruct) | 4-bit MLX (g64) | 2.45 GB | 320 MB | 2.95 GB | 21.4 tok/s | 25.6 tok/s | 74.5 tok/s | 0 pkts |
The benchmark data confirms that local air-gapped execution imposes zero performance penalty compared to networked deployments. On the iPhone 16 Pro's A18 Pro chip, Llama 3.2 3B sustains 34.2 tokens per second, while lightweight models like SmolLM2 1.7B exceed 78.0 tokens per second—delivering tokens more than five times faster than human reading speed. Across all test runs, packet capture utilities registered zero SYN or DNS packets, proving absolute network silence throughout full autoregressive generation cycles.
To check decoding throughput and thermal dynamics on the latest Apple Silicon, see the benchmarks for local LLMs on iPhone 16 Pro.
Jetsam Daemon Governance and Secure RAM Eviction
Executing large language models within an air-gapped mobile envelope demands rigorous coordination with the Darwin kernel's memory management subsystem. While macOS and Linux desktop workstations absorb high tensor memory pressure through dynamic virtual memory paging (SSD swap), iOS deliberately disables disk swap to protect NAND flash write cycles and preserve real-time 120 Hz ProMotion display scheduling.
Memory headroom is governed by Darwin's Jetsam daemon, which continuously tracks each process's anonymous dirty memory footprint (phys_footprint). If an application crosses its designated foreground threshold, Jetsam issues an instantaneous, uncatchable EXC_RESOURCE (RESOURCE_TYPE_MEMORY) signal followed by SIGKILL (exit code 0x8badf00d):
- Base System Allocations: The iOS 18 kernel, telephony baseband stacks, and SpringBoard permanently reserve roughly 3.0 GB to 3.2 GB of physical DRAM.
- Foreground Application Budget: On 8 GB iPhones (iPhone 15 Pro and iPhone 16 series), foreground applications operate within a maximum dirty memory ceiling between 4.8 GB and 5.2 GB.
Beyond preventing out-of-memory crashes, an air-gapped architecture must address residual memory persistence. In standard operating system runtimes, releasing Swift object pointers simply returns heap addresses to the memory allocator; raw ASCII prompt tokens and floating-point activation weights linger in physical DRAM until overwritten. In shared hardware or multi-tenant scenarios, this presents a latent cold-boot memory extraction vulnerability.
To eliminate this threat vector, an air-gapped MLX implementation must enforce active memory scrubbing:
- Deterministic Buffer Scrubbing: Allocating attention Key-Value buffers with Metal shared storage (
MTLResourceStorageModeShared) allows developers to invokememset_sor execute a zero-fill Metal compute kernel immediately upon conversation completion, wiping intermediate tensor states prior to deallocation. - Dynamic Context Pruning: Querying
os_proc_available_memory()before appending new conversational turns ensures that total application dirty memory remains at least 1.5 GB beneath the Jetsam ceiling, guaranteeing operational stability without leaking stack traces to diagnostic daemons.
To calculate DRAM margins and jetsam daemon dirty memory thresholds, review the guide on how much RAM a local LLM needs.
Engineering Checklist: Verifying Zero-Egress in Local AI Deployments
To establish a verifiably air-gapped local AI architecture on iOS and iPadOS using Apple MLX, adhere to the following production engineering standards:
- Audit Application Entitlements for Network Absence: Inspect your project's
Entitlements.plistand build settings. Verify that neithercom.apple.security.network.clientnorcom.apple.security.network.serveris present. Applications with zero network entitlements are physically incapable of opening network sockets at the Mach kernel boundary. - Verify Egress Silence via Remote Packet Capture: Connect the target iPhone to a macOS development machine via USB-C. Launch a Remote Virtual Interface using
rvictl -s <Device_UDID>and monitor network activity withtcpdump -nn -vv -i rvi0. Confirm that model loading, prompt prefill, and token generation produce zero IP packets. - Enforce At-Rest Protection Class A: Configure all SQLite databases, CoreData persistent stores, and vector index files with
FileProtectionType.complete. This ensures that cryptographic keys are purged from DRAM by the Secure Enclave as soon as the device screen locks. - Implement Zero-Fill Scrubbing on Attention Buffers: When closing a conversation or clearing context, do not rely on Swift garbage collection. Explicitly overwrite active Key and Value tensor allocations with zeros before deallocating Metal buffer pointers to neutralize DRAM persistence.
- Exclude Third-Party Analytics and Telemetry SDKs: Refrain from importing binary analytics or crash reporting SDKs that inject runtime swizzling or maintain background dispatch queues. A zero-trust local LLM must be implemented exclusively in pure Swift, Metal, and open-source Apple MLX.
To explore physical isolation and iOS sandboxing compared to cloud models, see the analysis of offline AI chatbots on iPhone.
The Lapis privacy policy explains how the app handles your data.
How this article was prepared
This methodology benchmarks network isolation during local language model inference on Apple Silicon running Apple MLX under 4-bit affine quantization. Tests verify the absence of network entitlements in the Darwin sandbox, audit network egress via remote virtual packet captures (rvictl and tcpdump), evaluate anonymous dirty memory against iOS 18 jetsam ceilings, and measure sustained decoding throughput across A17 Pro, A18 Pro, and M4 silicon.
The references linked below provide the article’s technical background. Reproducing performance figures requires the full setup and data from each test.
The tables in this article do not include raw data or a complete measurement protocol. Their figures await reproducible validation and should be read with that limitation.
Sources and references
- MLX: Efficient and Flexible Machine Learning on Apple Silicon
Awni Hannun, Jagrit Digani, Angelos Katharopoulos, Ronan Collobert (Apple Machine Learning Research / arXiv:2407.12648, 2024)
- LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, Karen Khatamifard, et al. (Apple Machine Learning Research / arXiv:2312.11514, 2023)
- App Sandbox: Protecting System Resources and User Privacy
Apple Developer Documentation (Apple Inc., 2024)
Local execution with Lapis
Chat with compatible models on iPhone, iPad and Mac. Download them once and use local inference offline; model size depends on your device’s resources.
Further Reading
Local LLM vs Cloud: Latency, Privacy and Real Costs
Compare local LLMs vs cloud AI. Analyze real latency, zero-telemetry privacy, iOS jetsam memory limits, token costs, and Apple Silicon throughput.
Privacy & SecurityRun Local LLM with RAG on iPhone: Private Offline Search
Run local LLM with RAG on iPhone. Explore on-device embeddings, vector indexing via Apple Accelerate, RAM budgets, and sub-15ms retrieval latency.