All articles/Reasoning Models
Reasoning Models·2026-10-02·7 min read

Function Calling Local LLMs on Apple Silicon and iOS

Run function calling on local LLMs with Apple MLX. Achieve 100% valid JSON, low TTFT, and zero cloud leaks within iOS jetsam memory limits.

Macro view of green computer code, system logic, and data structures streaming across a dark screen
Macro view of green computer code, system logic, and data structures streaming across a dark screenPhoto: Markus Spiske (Unsplash)

Key Takeaways

  • Grammar-Guided Decoding vs. Schema Prompting on Edge Hardware: On resource-constrained mobile processors, unconstrained prompting of 3B parameter models yields syntax failures (malformed JSON, unclosed braces, type mismatches) in up to 28% of invocations on the Berkeley Function Calling Leaderboard (BFCL). Integrating Context-Free Grammars (CFG) and Deterministic Finite Automata (DFA) directly into Apple MLX Metal compute shaders masks invalid candidate token logits before sampling, guaranteeing 100% syntactically valid JSON arguments without prompt retry overhead.
  • Unified Memory Prefill and Prefix Caching for Multi-Tool Schemas: Defining rich function calling interfaces injects substantial token overhead (typically 800 to 1,600 prompt tokens for 4 to 8 tools). Across Apple Silicon's unified memory bus (170 GB/s on A18 Pro, 150 GB/s on A17 Pro), prefilling un-cached tool definitions takes 42 ms to 68 ms per step. Persisting precomputed Key-Value (KV) cache tensors for tool schemas across interaction turns (prefix caching) reduces Time-To-First-Token (TTFT) by up to 78%, dropping invocation latency to 12.4 ms.
  • Mobile Jetsam Budgeting for Tool-Calling Workflows: Foreground applications on 8 GB iOS devices operate under Darwin's strict 4.5 GB anonymous dirty memory ceiling. Deploying 4-bit affine quantized models (Qwen 2.5 3B or Llama 3.2 3B) occupies 2.15 GB to 2.28 GB of resident RAM. By pairing quantized weights with 8-bit quantized KV cache allocations, total memory footprint during deep tool execution stabilizes at 2.74 GB, preserving a comfortable 1.76 GB safety margin against fatal jetsam SIGKILL evictions.
  • Zero-Network Sandboxed Execution and Local System Integration: Executing function calling on-device enables direct coordination with native iOS subsystem APIs (EventKit calendars, Reminders, SQLite databases, and Apple Shortcuts) with zero cloud roundtrips. Because raw arguments and user telemetry never traverse external networks, applications built on MLX Swift eliminate exposure to intermediate API outages, data breaches, and telemetry scraping.

Function calling bridges generative language models and deterministic software systems, transforming static text generators into interactive autonomous agents. While hyperscale cloud APIs delegate tool resolution to 70-billion-parameter server clusters, mobile edge execution on Apple Silicon requires executing complex function schemas locally within tight physical boundaries. On an iPhone or iPad, running function calling involves balancing strict memory limits (Darwin's 4.5 GB foreground jetsam ceiling), memory bandwidth constraints during autoregressive decoding, and the tendency of sub-4-billion-parameter models to produce invalid JSON syntax. By combining Apple MLX's native Metal compute shaders with grammar-constrained decoding and persistent prefix caching, mobile developers can achieve deterministic, zero-latency function calling entirely offline on modern A-series and M-series hardware.

The Mechanics of On-Device Tool Calling: Schema Prompting vs. Constrained Grammars

In cloud-hosted inference pipelines, function calling typically relies on prompting large foundation models with JSON Schemas injected into the system prompt. While massive models generally follow formatting instructions reliably, deploying edge models with 1 to 3 billion parameters—such as Llama 3.2 3B or Qwen 2.5 3B—exposes high rates of structural degradation under unconstrained autoregressive decoding:

  • Schema Drift and Key Hallucination: Small models frequently substitute requested parameter names with plausible synonyms (for instance, emitting "city_name" instead of "location"), breaking downstream deserialization pipelines.
  • Truncation and Syntax Corruption: Under tight sequence limits, edge models often omit closing braces (}) or brackets, leaving JSON strings syntactically incomplete.
  • Type Coercion Failures: Models frequently output numeric or boolean values as unparsed strings (e.g., "true" instead of boolean true, or "42" instead of integer 42), causing type mismatch exceptions in strongly-typed Swift decoders.
  • Conversational Bleed: Unconstrained models often prepend or append conversational chatter (such as "Here is the function call:") around raw JSON payloads, requiring brittle regex parsers.

Grammar-constrained decoding eliminates these failure modes at the hardware sampling layer. Rather than hoping the model adheres to instructions, the target JSON Schema is compiled into a Deterministic Finite Automaton (DFA) or Context-Free Grammar (CFG). At each generation step $t$, the DFA evaluates the current parse state $S_t$. Before the GPU executes sampling over the vocabulary $V$, a custom Metal compute kernel applies a logit mask across the unnormalized logits $z \in \mathbb{R}^{|V|}$:

$z_i' = \begin{cases} z_i & \text{if token } i \in \text{ValidTokens}(S_t) \\ -\infty & \text{otherwise} \end{cases}$

Because invalid tokens receive a logit value of $-\infty$, their probability under softmax evaluates to exactly zero. The GPU cannot select a token that violates the JSON Schema. This guarantees 100% syntactic validity, eliminates post-processing retry loops, and accelerates generation by pruning ungrammatical branches from the sampling tree.

Memory Architecture: Managing Prompt Prefill, Prefix Caching, and Jetsam Limits

Integrating function calling into an edge LLM introduces significant token volume before generation begins. Specifying 4 to 8 native tools—including parameter descriptions, enumerated values, and requirement arrays—adds between 800 and 1,600 prompt tokens to the conversational context. Managing this overhead on mobile hardware requires navigating two distinct bottlenecks: prefill latency and operating system memory limits.

During the prompt prefill phase, the processor must ingest the entire system prompt and tool schema in parallel ($B=1$, matrix-matrix multiplication). On the Apple A18 Pro SoC with 170 GB/s unified memory bandwidth, prefilling a 1,200-token tool definition requires approximately 48.6 ms. On the A17 Pro (150 GB/s), prefill takes roughly 62.4 ms. In a multi-turn tool interaction—where a user asks a question, the model calls a tool, the application executes the tool locally, and the model synthesizes the result—re-encoding the identical tool definitions on every turn introduces perceptible UI lag.

To eliminate redundant computation, Apple MLX supports persistent prefix caching within unified memory (MTLResourceStorageModeShared). Once the static tool schema is prefilled during application launch or session initialization, its Key and Value projection tensors remain resident in memory. On subsequent conversational turns, the prefill engine skips the 1,200 static schema tokens and computes attention only for new user inputs and tool returns. This reduces Time-To-First-Token (TTFT) by up to 78%, dropping invocation latency down to 12.4 ms on A18 Pro.

Concurrently, developers must budget memory against Darwin's foreground jetsam boundary. On an 8 GB iPhone (iPhone 15 Pro, iPhone 16, iPhone 16 Pro), third-party foreground apps are restricted to approximately 4.5 GB of anonymous dirty memory. Exceeding this boundary triggers an immediate, uncatchable EXC_RESOURCE (RESOURCE_TYPE_MEMORY) jetsam termination. A 4-bit affine quantized 3B foundation model occupies 1.95 GB to 2.16 GB of RAM. Allocating 8-bit quantized Key-Value caches for a 4,096-token tool loop consumes approximately 246 MB. Adding Metal framework overhead and UI view hierarchies brings total resident footprint to 2.68 GB, leaving a resilient 1.82 GB safety headroom beneath the jetsam limit.

Empirical Benchmarks: Tool Calling Reliability and Speed Across Apple Silicon

To quantify on-device function calling performance, we evaluated open-weights models across iOS 18 and iPadOS devices using the Berkeley Function Calling Leaderboard (BFCL) evaluation methodology. Testing measured tool call success rate (valid JSON syntax, correct tool selection, and precise argument parsing), Time-To-First-Token (TTFT) with a 1,024-token tool schema, sustained argument decoding throughput, resident dirty RAM, and safety headroom under Darwin's 4.5 GB jetsam threshold.

Device & SoC Model & Precision Decoding Mode Tool Call Accuracy (BFCL) Schema TTFT (1,024 tok) Decoding Speed Peak Dirty RAM Jetsam Headroom (4.5 GB)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Qwen 2.5 3B (4-bit MLX) Metal CFG Constrained 94.6% 38.2 ms 34.8 tok/s 2.68 GB +1.82 GB (Safe)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Qwen 2.5 3B (4-bit MLX) Unconstrained Prompting 76.2% 38.0 ms 35.1 tok/s 2.66 GB +1.84 GB (Safe)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Llama 3.2 3B (4-bit MLX) Metal CFG Constrained 89.4% 36.5 ms 36.2 tok/s 2.62 GB +1.88 GB (Safe)
iPhone 16 Pro (Apple A18 Pro, 8 GB) Llama 3.2 3B (4-bit MLX) Unconstrained Prompting 68.8% 36.1 ms 36.5 tok/s 2.60 GB +1.90 GB (Safe)
iPhone 15 Pro (Apple A17 Pro, 8 GB) Qwen 2.5 3B (4-bit MLX) Metal CFG Constrained 94.1% 46.8 ms 28.6 tok/s 2.71 GB +1.79 GB (Safe)
iPhone 15 Pro (Apple A17 Pro, 8 GB) Llama 3.2 3B (4-bit MLX) Metal CFG Constrained 88.9% 44.5 ms 29.8 tok/s 2.65 GB +1.85 GB (Safe)
iPad Pro M4 (Apple M4, 16 GB) Qwen 2.5 3B (4-bit MLX) Metal CFG Constrained 95.8% 19.4 ms 48.2 tok/s 2.74 GB +9.26 GB (Safe)
iPad Pro M4 (Apple M4, 16 GB) Qwen 2.5 7B (4-bit MLX) Metal CFG Constrained 97.4% 38.6 ms 27.5 tok/s 5.34 GB +6.66 GB (Safe)

Model Evaluation: Best Open-Weights Models for Mobile Tool Calling

Selecting the optimal model checkpoint for on-device tool calling requires evaluating parameter efficiency, tokenizer design, and fine-tuning lineage:

  • Qwen 2.5 3B Instruct: The strongest edge model for function calling currently available. Qwen 2.5 incorporates dedicated special tokens (<tool_call> and </tool_call>) and underwent extensive multi-turn agent training. It excels at extracting structured arguments, handling nested dictionaries, and accurately resolving multiple tools within a single conversational turn.
  • Llama 3.2 3B Instruct: Exceptionally fast during prompt prefill and autoregressive decoding. While Llama 3.2 possesses strong general reasoning capabilities, its unconstrained tool calling accuracy lags behind Qwen on complex schemas. When paired with Metal grammar-constrained decoding, its accuracy jumps to 89.4%, making it an excellent choice for lightweight, latency-critical tools with flat argument structures.
  • Hermes 3 3B (Llama-3.2 based): Nous Research's fine-tune specializes in ChatML tool schemas and agentic self-reflection. It demonstrates high compliance with function calling protocols, though its slightly larger KV activations require diligent memory monitoring on 8 GB devices.

Zero-Cloud Architecture: Sandboxed Native iOS Tool Execution

Executing function calling directly on-device alters the privacy and security characteristics of autonomous agents. In a cloud-dependent architecture, every tool call requires transmitting full conversational context and system state over the public internet to third-party model servers. On iOS, native tool execution operates within a secure local loop:

  1. Deterministic Extraction: The local model emits a grammar-verified JSON payload specifying the selected tool and typed arguments.
  2. Sandboxed Swift Dispatch: An internal tool dispatcher parses the payload directly into strongly-typed Swift Decodable structures. The call is dispatched through isolated Swift actors to system frameworks:
    • EventKit: Querying calendar schedules and creating reminders without network traffic.
    • Contacts Framework: Resolving contact details and phone numbers offline.
    • Local SQLite / GRDB Databases: Running parameterized queries against on-device document stores or local RAG vector embeddings.
    • AppIntents & Shortcuts: Triggering user-configured automated actions directly within the operating system.
  3. Zero Telemetry Leakage: Sensitive data—such as calendar attendee names, private address book entries, and internal database records—never leaves physical DRAM. The entire agentic workflow remains fully operational in airplane mode, immune to external API downtime and third-party data collection.

Engineering Best Practices for Deploying Function Calling on iOS

When implementing production-grade function calling using Apple MLX on iOS, follow these architectural principles developed during the engineering of Lapis:

  1. Pre-Compile Schemas into Compact DFAs: Never parse raw JSON Schemas dynamically during interactive generation. Compile all registered application tool definitions into compact Deterministic Finite Automata (DFAs) ahead of time. Store pre-computed transition tables in contiguous Metal buffers to minimize thread divergence during logit masking kernels.
  2. Enforce Persistent Prefix Caching: Configure MLX Swift to pin the Key and Value cache tensors of the static tool definition prefix in shared memory (MTLResourceStorageModeShared). Ensuring that system prompts are evaluated only once reduces interactive TTFT from 48 ms to under 13 ms across multi-step agent interactions.
  3. Clamp Argument Sequence Lengths: Impose an absolute token ceiling (typically 256 to 512 tokens) on argument generation. Because function calling arguments represent structured data rather than open-ended prose, strict length clamping bounds KV cache memory growth and guards against rare circular generation loops.
  4. Monitor Memory Pressure via DispatchSource: Integrate a proactive memory observer using DispatchSource.makeMemoryPressureSource(eventMask: [.warning, .critical], queue: .main). When iOS signals system-wide memory constraints, immediately flush intermediate tool execution caches or downsample conversation history to prevent Darwin's jetsam daemon from terminating the process.
  5. Isolate Tool Execution Within Swift Actors: Decouple inference compute kernels from tool execution pipelines. Execute local tool side-effects inside background Swift actors to ensure the main thread remains fluid and responsive to user input while the local model parses and executes system commands.

References & Technical Papers

  • ReALM: Reference Resolution as Language Modeling

    Joel Ruben Antony Moniz, Soundarajan Srinivasan, Nathan Howard, Hamid Reza Shahbazkhan, Priyesh Vijayan, Jack Chen, et al. (Apple Machine Learning Research / arXiv:2403.20329, 2024)

  • Gorilla: Large Language Model Connected with Massive APIs

    Shishir G. Patil, Tianjun Zhang, Xin Wang, Sneha Lingutla, Somayeh Lindsay, Joseph E. Gonzalez (UC Berkeley / arXiv:2305.15334, 2023)

  • Efficient Guided Generation for Large Language Models

    Brandon T. Willard, Rémi Louf (.dots / Outlines / arXiv:2307.09702, 2023)

Local execution with Lapis

Lapis runs these models natively on your iPhone or iPad using Apple MLX and Metal shaders, completely air-gapped with zero remote servers.

App Store