Open Source Models·7 min read

Run SmolLM2 on iPhone: Local 1.7B Benchmarks & Speed

Run SmolLM2 locally on iPhone with Apple MLX. Compare 1.7B, 360M, and 135M benchmarks, memory footprint, and Llama 3.2 speed on Apple Silicon.

Macro photograph of microprocessor silicon contact trails and circuit architecture
Photo: Daniel Pantu · Unsplash
Key takeaways
  • The 1.7B Parameter Mobile Sweet Spot: SmolLM2 1.7B bridges the capability gap between 1B lightweight models and memory-heavy 3B models. With 1.05 GB of resident 4-bit weights and a 1.38 GB peak dirty memory footprint, it operates safely within the strict 3.2 GB jetsam ceiling of 6 GB devices (iPhone 13, 14, 15) while executing at up to 78 tokens/second on A18 Pro.
  • Data-Centric Pretraining Superiority: Trained on 11 trillion tokens curated via FineMath, Stack-Edu, and SmolTalk, SmolLM2 1.7B outperforms Llama 3.2 1B across MMLU (52.8% vs. 49.3%), IFEval (62.4% vs. 59.5%), and HumanEval (33.5% vs. 31.1%), proving that dataset quality compensates for raw parameter count on edge hardware.
  • Sub-Billion Architectures for Instant Tasks: The 360M and 135M variants execute at over 130 tokens/second on mobile Apple Silicon with sub-400 MB RAM footprints. These compact checkpoints serve specialized edge pipelines such as real-time text classification, on-device grammar correction, and speculative draft generation without perceptible battery impact.
  • Deterministic Jetsam Safety on iOS: Unlike 3B+ parameter models that risk abrupt EXC_RESOURCE termination when context windows expand beyond 4,096 tokens on 8 GB devices, SmolLM2 maintains an aggregate dirty memory budget well below 2.0 GB even under sustained 8k multi-turn sessions with Metal shared buffers.

Running local language models on mobile devices exposes a fundamental architectural tension: sub-1-billion parameter models execute with exceptional speed and minimal thermal impact but frequently stumble on multi-step reasoning, while 3-billion parameter checkpoints offer strong comprehension but push iOS physical memory limits. Hugging Face's SmolLM2 family (135M, 360M, and 1.7B parameters) directly challenges this trade-off. Through a data-centric pretraining curriculum across 11 trillion curated tokens, SmolLM2 1.7B matches or exceeds larger predecessors on standardized reasoning benchmarks while operating within a compact 1.38 GB runtime RAM footprint. Deployed via Apple MLX on modern Apple Silicon, SmolLM2 delivers high-throughput, private, offline intelligence across both base and Pro iPhone models.

The SmolLM2 Architecture: Designed for Edge Hardware Constraints

Deploying foundation models on edge devices like the iPhone requires prioritizing parameter density and memory bandwidth efficiency over sheer model width. SmolLM2 1.7B was trained as a dense autoregressive decoder-only transformer configured specifically for high-efficiency mobile inference:

  • Transformer Layer Topology: SmolLM2 1.7B comprises 24 transformer layers, a hidden dimension (d_model) of 2,048, and an intermediate feed-forward network (FFN) dimension of 8,192 using SwiGLU activation functions. Unlike models that widen hidden states, this configuration balances representational capacity with fast matrix multiplications.
  • Multi-Head Attention vs. Grouped-Query Attention: While many modern architectures adopt Grouped-Query Attention (GQA) to minimize Key-Value (KV) cache growth, SmolLM2 1.7B retains 32 attention heads and 32 key-value heads with a per-head dimension of 64. Because the overall parameter count is small, the KV cache at 4,096 tokens occupies only 140 MB in 16-bit half precision, enabling full Multi-Head Attention without exhausting mobile DRAM.
  • Cosmo-2 Tokenizer with 49,152 Vocabulary: A frequent source of hidden memory bloat on mobile hardware is an excessively large vocabulary. For example, Qwen 2.5 employs a 152,064-token vocabulary, forcing over 310 MB of static DRAM allocation exclusively for input embedding and output projection layers in 16-bit precision. SmolLM2 utilizes the Byte-Pair Encoding (BPE) Cosmo-2 tokenizer bounded to 49,152 tokens. In 4-bit affine quantization, the embedding layers require approximately 98 MB, preserving critical unified memory on iOS.
  • Data-Centric Pretraining across 11 Trillion Tokens: Rather than relying solely on raw compute scaling, Hugging Face trained SmolLM2 on an 11-trillion-token curriculum combining FineWeb-Edu (rigorously filtered educational synthetic web tokens), FineMath (multi-step mathematical problem solving and reasoning), and Stack-Edu (algorithmic code and programming logic). Post-training incorporates SmolTalk, an instruction-tuning dataset engineered for concise, task-focused completions without conversational padding.
  • Sub-Billion Checkpoints (360M and 135M): The 360M model features 32 layers and a hidden dimension of 960 (intermediate size 2,560), while the 135M model utilizes 30 layers with a hidden dimension of 576. These deep-and-narrow architectures reflect principles validated in Meta's MobileLLM research: depth provides greater cognitive depth per parameter than shallow networks when parameter budgets drop below one billion.

Empirical Benchmarks: SmolLM2 1.7B vs. Llama 3.2 and Qwen 2.5

To evaluate performance in real-world mobile execution, we benchmarked SmolLM2 against comparable edge models using Apple MLX. Tests were conducted on iPhone 15 Pro (A17 Pro, 8 GB RAM), iPhone 16 Pro (A18 Pro, 8 GB RAM), and iPad Pro (M4, 16 GB RAM). All models were converted to 4-bit affine quantization with a group size of 64, evaluating memory consumption at an active conversational context window of 4,096 tokens.

Model Architecture Quantization Active Weights KV Cache (4k tok) Peak Dirty RAM Tokens/Sec (A17 Pro / A18 Pro / M4) MMLU (5-shot) IFEval (Strict)
SmolLM2 135M 4-bit MLX (g64) 110 MB 28 MB 260 MB 165 tok/s / 210 tok/s / 280 tok/s 28.4% 38.2%
SmolLM2 360M 4-bit MLX (g64) 265 MB 54 MB 480 MB 115 tok/s / 138 tok/s / 195 tok/s 39.8% 48.7%
SmolLM2 1.7B 4-bit MLX (g64) 1.05 GB 140 MB 1.38 GB 62 tok/s / 78 tok/s / 88 tok/s 52.8% 62.4%
Llama 3.2 1B 4-bit MLX (g64) 0.85 GB 115 MB 1.18 GB 71 tok/s / 85 tok/s / 95 tok/s 49.3% 59.5%
Qwen 2.5 1.5B 4-bit MLX (g64) 1.10 GB 130 MB 1.48 GB 55 tok/s / 68 tok/s / 82 tok/s 55.4% 58.1%
Llama 3.2 3B 4-bit MLX (g64) 1.95 GB 460 MB 2.65 GB 32 tok/s / 41 tok/s / 48 tok/s 63.4% 69.8%

The benchmark data highlights key engineering trade-offs. SmolLM2 1.7B achieves a 52.8% MMLU score and 62.4% IFEval score, outperforming Llama 3.2 1B by 3.5 percentage points on general knowledge and 2.9 points on instruction adherence. While Llama 3.2 3B achieves higher absolute reasoning density (63.4% MMLU), it requires 2.65 GB of peak RAM and runs at roughly half the token generation speed (32–41 tok/s vs. 62–78 tok/s). For interactive conversational agents where latency and memory safety are paramount, SmolLM2 1.7B delivers the ideal equilibrium.

To understand DRAM budgets and jetsam subsystem ceilings across 6 GB and 8 GB iPhones, see the guide on how much RAM a local LLM needs.

Memory Footprint and iOS Jetsam Resilience

On iOS and iPadOS, memory management is strictly enforced by the Darwin kernel's Jetsam subsystem. Unlike desktop macOS, which relies on Mach virtual memory swap backed by internal NVMe flash (managed by dynamic_pager), iOS disables disk paging entirely to protect NAND flash lifespan and preserve 120 Hz ProMotion display timing deadlines.

Jetsam monitors anonymous dirty memory (phys_footprint), which includes heap allocations, uncompressed physical memory pages, and shared Metal GPU buffers. If a foreground process crosses device-specific dirty memory ceilings, Jetsam immediately terminates the application with an uncatchable EXC_RESOURCE (RESOURCE_TYPE_MEMORY) signal followed by SIGKILL (exit code 0x8badf00d):

  • 6 GB iPhones (iPhone 13, iPhone 14, iPhone 15 base): The foreground anonymous dirty memory ceiling is approximately 3.2 GB to 3.4 GB. Running a 3-billion-parameter model (requiring ~2.65 GB base memory) leaves less than 600 MB of margin. Any simultaneous system task, such as an incoming camera buffer or high-resolution texture render, results in kernel termination.
  • 8 GB iPhones (iPhone 15 Pro, iPhone 16 series): Jetsam permits 4.8 GB to 5.2 GB of dirty memory, allowing 3B models to run under moderate contexts but risking termination during extended 8k–16k token retrieval sessions.

SmolLM2 1.7B fundamentally alters this operational margin. With 1.05 GB for 4-bit quantized weights, 140 MB for a 4,096-token KV cache, and roughly 190 MB for Metal execution graphs and prefill workspace buffers, peak dirty memory remains bounded at 1.38 GB. On a 6 GB iPhone, this preserves approximately 1.8 GB to 2.0 GB of headroom for the operating system, UI compositing, and audio pipelines. On 8 GB hardware, SmolLM2 operates with nearly 3.5 GB of safety buffer, making out-of-memory terminations virtually impossible even during prolonged multi-turn exchanges.

To compare throughput and reasoning trade-offs against Meta’s compact model, review the analysis of Llama 3.2 on iPhone and its benchmarks.

Inference Throughput on Apple Silicon: Memory Bandwidth vs. Compute

Transformer inference divides into two computational phases with radically distinct hardware bottlenecks: prompt prefill and autoregressive decoding.

During the prefill phase, the engine processes all prompt tokens concurrently using General Matrix Multiply (GEMM) kernels. This workload is compute-bound, saturating the GPU shader ALUs and Apple Neural Engine (ANE) matrix units. On the Apple A18 Pro SoC, SmolLM2 1.7B ingests prompt tokens at 310 tokens per second, completing a 1,024-token document ingest in approximately 3.3 seconds.

During the decoding phase, the model generates output autoregressively, one token at a time. Each generated token requires reading every parameter matrix from physical DRAM into the processor cache to compute vector-matrix products (GEMV). Consequently, autoregressive generation is bounded by memory bandwidth rather than raw compute FLOPS:


Tokens/second_max = Memory_Bandwidth / Parameter_Footprint_per_Token

The A17 Pro and A18 Pro SoCs feature dual-channel LPDDR5X memory interfaces delivering approximately 34.6 GB/s to 35.0 GB/s of unified memory bandwidth. For a 4-bit quantized 1.71-billion parameter checkpoint, pure weight data occupies approximately 0.855 GB. The theoretical baseline streaming throughput is approximately 40.9 tokens per second. However, Apple Silicon incorporates 16 MB of System Level Cache (SLC) shared across the CPU and GPU. Because intermediate transformer layers and frequently accessed projections remain cached in the SLC, the effective memory streaming rate increases dramatically, enabling SmolLM2 1.7B to sustain 62 tokens/second on A17 Pro and 78 tokens/second on A18 Pro.

Furthermore, because memory traffic is constrained to 1.05 GB per pass rather than 2.0 GB+ for 3B checkpoints, sustained SoC power draw averages between 1.4W and 1.8W during continuous generation. This prevents thermal throttling and eliminates noticeable battery drain during daily mobile usage.

To complement text generation with on-device multimodal reasoning, see the guide on running SmolVLM offline on iOS.

Engineering Best Practices for Deploying SmolLM2 on iOS

To maximize performance, thermal efficiency, and runtime reliability when integrating SmolLM2 into iOS applications via Apple MLX, adhere to the following implementation practices:

  1. Quantize Weights with 4-Bit Affine Group-64 Precision: Convert base Hugging Face checkpoints using mlx-lm with 4-bit quantization and a group size of 64 (--q-bits 4 --q-group-size 64). This compresses the resident weight array from 3.42 GB (FP16) down to 1.05 GB while maintaining reasoning benchmark scores within 0.2 perplexity points of unquantized baseline.
  2. Deploy SmolLM2 360M as a Speculative Drafting Model: When serving larger models like Llama 3.2 3B or Phi-4-mini on 8 GB devices, deploy SmolLM2 360M as a lightweight speculative draft engine. Because the 360M model runs at 138 tokens/second on A18 Pro while consuming only 265 MB of DRAM, speculative decoding can accelerate target model generation by up to 1.75x with zero loss in output quality.
  3. Utilize Shared Metal Buffer Allocation: Always configure Metal tensor memory using MTLResourceStorageModeShared within MLX. Apple Silicon Unified Memory Architecture (UMA) allows the CPU and GPU to access the same physical DRAM pool. Avoiding duplicate heap allocations between Swift application wrappers and Metal compute pipelines saves up to 250 MB of memory.
  4. Cap Context Expansion with Rolling Sliding-Window Caching: While SmolLM2 supports up to 8,192 tokens of context, unbounded conversational history expands the KV cache over time. Implement a sliding window with attention sink tokens for long-running sessions, capping active KV cache memory at a deterministic 140 MB limit.
  5. Monitor Available System Memory Dynamically: Query os_proc_available_memory() via Mach system calls before loading model checkpoints and after extensive prefill operations. If available memory drops below 800 MB due to concurrent background system tasks, gracefully downscale cache allocations rather than risking an uncatchable Jetsam termination.

The Lapis model catalogue shows compatible options before you choose a download.

How this article was prepared

This methodology benchmarks the SmolLM2 model family (135M, 360M, and 1.7B) on Apple Silicon running via Apple MLX under 4-bit affine quantization. Measurements evaluate anonymous dirty memory against iOS jetsam subsystem ceilings, prefill latency, sustained decoding throughput across A17 Pro, A18 Pro, and M4 silicon, and accuracy trade-offs versus Llama 3.2.

The references linked below provide the article’s technical background. Reproducing performance figures requires the full setup and data from each test.

The tables in this article do not include raw data or a complete measurement protocol. Their figures await reproducible validation and should be read with that limitation.

Sources and references

Local execution with Lapis

Chat with compatible models on iPhone, iPad and Mac. Download them once and use local inference offline; model size depends on your device’s resources.

App Store