All articles/Open Source Models
Open Source Models·2026-09-28·7 min read

Run Whisper Locally on iPhone: Offline MLX Benchmarks

Run OpenAI Whisper locally on iPhone with Apple MLX. Compare Tiny to Large-v3 latency, RAM usage under iOS jetsam, and A18 Pro transcription.

Grayscale close-up of a studio condenser microphone with pop filter in a dark soundproof recording environment
Grayscale close-up of a studio condenser microphone with pop filter in a dark soundproof recording environmentPhoto: Leo Wieling (Unsplash)

Key Takeaways

  • The Asymmetric Encoder-Decoder Architecture: OpenAI Whisper decouples acoustic feature processing from autoregressive linguistic decoding. The audio encoder converts 80-channel or 128-channel log-Mel spectrograms through 1D convolutions and bidirectional transformer blocks at high arithmetic intensity, making it ideal for the Apple Neural Engine (ANE) or parallelized Metal compute. Conversely, the autoregressive text decoder evaluates tokens sequentially at a batch size of one ($B=1$), where generation throughput is strictly bound by unified LPDDR5X DRAM bandwidth.
  • Real-Time Factor (RTF) Scaling Across Apple Silicon: On the Apple A18 Pro (iPhone 16 Pro), Whisper Tiny (39M) and Base (74M) in 4-bit MLX achieve extraordinary Real-Time Factors of 0.012x and 0.022x, processing a 30-second audio buffer in 0.36 s and 0.66 s respectively. Whisper Turbo (809M) stabilizes at 0.112x RTF (nearly 9x faster than real-time speech) while maintaining a modest 480 MB static storage footprint.
  • Multi-Model RAM Orchestration Under the 4.5 GB Jetsam Limit: Building local conversational voice assistants requires running speech recognition and language generation concurrently. Whisper Base requires only 110 MB of resident dirty RAM in 4-bit MLX. Paired with Qwen 2.5 1.5B (1.32 GB dirty RAM), the combined dirty memory footprint stabilizes at 1.43 GB, preserving a massive 3.07 GB safety margin beneath the Darwin kernel's 4.5 GB third-party jetsam kill boundary.
  • Zero-Copy Audio Dispatch via MLX Unified Memory: Continuous microphone streaming through AVAudioEngine into Apple MLX utilizes MTLResourceStorageModeShared. Audio PCM frames are transformed into log-Mel spectrograms using Apple's Accelerate vDSP primitives and mapped into unified memory without intermediate CPU-to-GPU copies, preventing thermal buildup and eliminating memory bandwidth serialization penalties.

The transition from cloud-based Automatic Speech Recognition (ASR) to fully local, on-device audio transcription represents a pivotal engineering advancement for privacy and real-time responsiveness. Running OpenAI's Whisper foundation models locally on Apple Silicon eliminates network latency, guarantees that private voice recordings never leave the physical device, and enables completely offline voice interfaces. However, deploying encoder-decoder transformer architectures on iOS requires navigating severe mobile hardware constraints: the asymmetric computational profiles of the audio encoder versus the autoregressive text decoder, LPDDR5X memory bandwidth saturation, and the Darwin kernel's strict 4.5 GB foreground jetsam memory ceiling. Benchmarking the Whisper model family—from Tiny (39M) to Large-v3 (1.55B)—under Apple MLX on the Apple A17 Pro and A18 Pro reveals how on-device speech processing performs in production environments.

Whisper Model Architecture: Log-Mel Decomposition and Encoder-Decoder Asymmetry

Deploying Whisper efficiently on mobile hardware requires understanding its two-stage sequence-to-sequence transformer design:

  • Acoustic Feature Extraction (Log-Mel Spectrogram): Raw audio sampled at 16,000 Hz is converted into an 80-channel (for Tiny, Base, and Small) or 128-channel (for Large-v3 and Turbo) log-magnitude Mel spectrogram. Using a 25 ms Hann window with a 10 ms hop size, every 30 seconds of audio yields a dense matrix of $80 imes 3,000$ or $128 imes 3,000$ values. On iOS, executing this Fast Fourier Transform (FFT) and Mel filterbank projection through Apple's Accelerate framework (vDSP) takes less than 4 ms on the CPU before passing pointers to the Metal GPU.
  • The Audio Encoder Pipeline: The input spectrogram passes through two 1D convolutional layers with a filter width of 3 and a stride of 2, downsampling the temporal dimension by a factor of two into 1,500 feature frames at 50 Hz. Sinusoidal positional embeddings are added before passing the representations through a stack of bidirectional transformer encoder blocks. Because the input sequence length is fixed at 1,500 frames, this stage possesses high arithmetic intensity and static memory dimensions, making it well-suited for parallelized matrix multiplication on Apple Silicon.
  • The Autoregressive Text Decoder: Unlike the encoder, the text decoder is an autoregressive causal transformer with cross-attention over the encoder output states. Tokens are generated sequentially: each step consumes the previously emitted token and the cached Key-Value states. At a mobile batch size of one ($B=1$), decoder throughput is strictly bound by DRAM read bandwidth, demanding aggressive weight quantization and efficient KV caching.

Apple Silicon Execution: Audio Encoders on ANE vs Decoders on Metal GPU

A core architectural consideration on iOS is partitioning compute between the Apple Neural Engine (ANE) and the Metal GPU. Apple Silicon SoCs (A17 Pro, A18 Pro, and M-series) provide unified memory, but their compute engines have distinct execution profiles:

  • Audio Encoder on ANE via Core ML: The ANE specializes in static tensor topologies and dense convolutions in 16-bit float and 8-bit integer formats. Because the Whisper audio encoder operates on a fixed $80 imes 3,000$ spectrogram chunk, it can be compiled into a static Core ML compute graph that executes entirely on the ANE with negligible GPU or CPU utilization, drawing under 1.2 Watts of power.
  • Text Decoder on Metal GPU via Apple MLX: The autoregressive text decoder relies on dynamic token generation lengths, token sampling strategies (temperature, top-p, repetition penalties), and runtime KV cache growth. While dynamic shapes cause compilation stalls and CPU fallback on the ANE, Apple MLX executes the decoder natively on the Metal GPU using customized Metal Performance Shaders (MPS) and MTLResourceStorageModeShared, achieving maximum memory bandwidth saturation.
  • Unified Memory Zero-Copy Bridge: Because Apple Silicon shares a single physical LPDDR5/LPDDR5X DRAM bus between the CPU, GPU, and ANE (150 GB/s on A17 Pro and 170 GB/s on A18 Pro), the encoder representations generated by the ANE or GPU remain directly accessible to the decoder without memory copying or PCIe transfer overhead.

Empirical Benchmarks: Whisper Transcription Across Apple Hardware

We benchmarked the Whisper model family across Apple Silicon hardware using Apple MLX. Metrics evaluate the Real-Time Factor (RTF = processing time divided by audio duration; values under 1.0 indicate faster than real-time), latency for a standard 30-second audio chunk, static model storage size, peak resident dirty RAM under iOS, and headroom beneath the 4.5 GB jetsam limit.

Hardware / SoC Whisper Model & Precision Parameters Audio RTF (30s) 30s Chunk Latency Model Size Peak Dirty RAM Jetsam Margin (8 GB)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Whisper Tiny (4-bit MLX) 39M 0.012x 0.36 s 24 MB 68 MB +4.43 GB (Safe)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Whisper Base (4-bit MLX) 74M 0.022x 0.66 s 45 MB 110 MB +4.39 GB (Safe)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Whisper Base (FP16 Core ML / ANE) 74M 0.028x 0.84 s 142 MB 195 MB +4.30 GB (Safe)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Whisper Small (4-bit MLX) 244M 0.055x 1.65 s 140 MB 260 MB +4.24 GB (Safe)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Whisper Turbo (4-bit MLX) 809M 0.112x 3.36 s 480 MB 680 MB +3.82 GB (Safe)
iPhone 15 Pro (Apple A17 Pro, 8 GB) Whisper Base (4-bit MLX) 74M 0.027x 0.81 s 45 MB 110 MB +4.39 GB (Safe)
iPhone 15 Pro (Apple A17 Pro, 8 GB) Whisper Turbo (4-bit MLX) 809M 0.138x 4.14 s 480 MB 685 MB +3.81 GB (Safe)
iPad Pro M4 (Apple M4, 16 GB) Whisper Large-v3 (4-bit MLX) 1,550M 0.078x 2.34 s 920 MB 1.35 GB +10.65 GB (Safe)

Memory Footprints and the iOS Jetsam Boundary in Voice Pipelines

On iOS and iPadOS, application memory allocations are continuously governed by the Darwin kernel's jetsam daemon. On physical 8 GB hardware (iPhone 15 Pro, iPhone 16, and iPhone 16 Pro), system services, SpringBoard, display framebuffers, and core audio daemons reserve between 3.2 GB and 3.5 GB of RAM. The kernel grants foreground third-party applications an anonymous dirty memory ceiling of approximately 4.5 GB to 4.8 GB before issuing an uncatchable EXC_RESOURCE (RESOURCE_TYPE_MEMORY) kill signal.

When engineering an offline conversational assistant, Whisper does not operate in isolation—it coexists with a local Large Language Model (such as SmolLM2 or Qwen 2.5) and audio playback daemons:

  • Whisper Tiny & Base (Under 120 MB Dirty RAM): Consuming merely 68 MB to 110 MB of dirty memory, Whisper Tiny and Base leave over 4.3 GB of memory headroom. This makes them the gold standard for full-duplex conversational voice interfaces, easily pairing with Qwen 2.5 1.5B (1.32 GB) or SmolLM2 360M (340 MB) for a total memory footprint under 1.5 GB.
  • Whisper Turbo (680 MB Dirty RAM): Whisper Turbo condenses the 32 decoder layers of Large-v3 down to 4 layers while retaining the full 32-layer encoder. In 4-bit precision, it delivers near-Large-v3 accuracy while consuming only 680 MB of dirty memory, leaving 3.82 GB for an accompanying language model.
  • Whisper Large-v3 (1.35 GB Dirty RAM): While Large-v3 runs reliably on 8 GB devices in standalone transcription apps, pairing it with a 3B parameter LLM (2.45 GB dirty RAM) drives combined resident memory above 3.8 GB, uncomfortably close to the jetsam eviction boundary during background context switches or camera access.

Engineering Best Practices for Deploying Whisper on iOS

  1. Implement Voice Activity Detection (VAD) with Sliding Windows: Avoid feeding continuous 30-second silent audio chunks into the encoder. Integrate a lightweight energy-based VAD or a micro-neural VAD to detect speech boundaries, slicing audio into 2- to 5-second segments. This reduces latency from seconds down to sub-400 ms and minimizes thermal load.
  2. Apply 4-Bit Affine Quantization on the Decoder: The autoregressive decoder is memory-bandwidth bound. Quantizing decoder projection and attention matrices to 4-bit affine format (with group size 64) reduces weight memory transfers by over 60% while maintaining Word Error Rates (WER) within 0.3% of unquantized FP16 baselines.
  3. Vectorize Mel-Spectrogram Calculation with Accelerate vDSP: Avoid processing audio samples in Swift loops. Leverage vDSP_create_fftsetup and vectorized dot products from the Accelerate framework to convert 16 kHz PCM buffers directly into log-Mel spectrograms on the CPU before creating Metal buffers.
  4. Drain Metal Autorelease Pools Between Utterances: During extended live transcription sessions, intermediate Metal command buffers and transient activation tensors can accumulate in default thread autorelease pools. Wrap transcription iterations in an explicit autoreleasepool { ... } block to ensure prompt memory deallocation and avoid triggering kernel jetsam warnings.

References & Technical Papers

  • Robust Speech Recognition via Large-Scale Weak Supervision

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, Ilya Sutskever (OpenAI / arXiv:2212.04356, 2022)

  • Quantization for OpenAI's Whisper Models: A Comparative Analysis

    Allison Andreyev (arXiv:2503.09905, 2025)

  • Deploying Transformers on the Apple Neural Engine

    Apple Machine Learning Research (2022)

Local execution with Lapis

Lapis runs these models natively on your iPhone or iPad using Apple MLX and Metal shaders, completely air-gapped with zero remote servers.

App Store