On-device inference, privacy, and hardware.
Technical benchmarks, memory bandwidth analyses, and notes on running foundation models directly on Apple Silicon.
Function Calling Local LLMs on Apple Silicon and iOS
Run function calling on local LLMs with Apple MLX. Achieve 100% valid JSON, low TTFT, and zero cloud leaks within iOS jetsam memory limits.
Apple MLX vs Core ML: Which Runs Local LLMs Faster?
Compare Apple MLX and Core ML for local LLM inference on iOS. Analyze ANE limits, dynamic KV cache, memory bandwidth, and token speeds.
KV Cache Quantization: Fast Local LLMs on Apple Silicon
Learn how 4-bit and 8-bit KV cache quantization cuts RAM by 73% in Apple MLX, prevents iOS jetsam crashes, and boosts decoding on A18 Pro.
Run Gemma on iPhone: Local MLX Benchmarks & Speed
Run Google Gemma on iPhone with Apple MLX. Compare Gemma 2 2B latency, sliding window KV cache, RAM under iOS jetsam, and A18 Pro speed.
Run Whisper Locally on iPhone: Offline MLX Benchmarks
Run OpenAI Whisper locally on iPhone with Apple MLX. Compare Tiny to Large-v3 latency, RAM usage under iOS jetsam, and A18 Pro transcription.
Run Qwen on iPhone: Local MLX Benchmarks & Memory
Run Qwen 2.5 locally on iPhone with Apple MLX. Compare 0.5B to 7B benchmarks, 4-bit memory footprints, and A18 Pro tokens/s under iOS jetsam.
Speculative Decoding on Apple Silicon: MLX Speedup & Limits
Accelerate local LLM inference on Apple Silicon using MLX speculative decoding. Benchmark draft models, memory bandwidth, tokens/s, and iOS limits.
Apple Neural Engine vs GPU for LLMs: Latency & Memory
Compare Apple Neural Engine vs GPU for local LLM inference. Analyze unified memory bandwidth, Core ML vs MLX, ANE constraints, and mobile tokens/s.
Local LLM vs Cloud: Latency, Privacy and Real Costs
Compare local LLMs vs cloud AI. Analyze real latency, zero-telemetry privacy, iOS jetsam memory limits, token costs, and Apple Silicon throughput.
Run Local LLM with RAG on iPhone: Private Offline Search
Run local LLM with RAG on iPhone. Explore on-device embeddings, vector indexing via Apple Accelerate, RAM budgets, and sub-15ms retrieval latency.
Apple MLX vs llama.cpp: Mobile LLM Benchmarks
Compare Apple MLX and llama.cpp on Apple Silicon. Discover Metal kernel execution, KV cache RAM limits, prefill speed, and decode benchmarks.
Best Local AI App for iPhone: Architecture Guide
Discover the best local AI app for iPhone. Compare Apple MLX vs llama.cpp, RAM jetsam limits, 4-bit quantization, and zero-cloud privacy.
Run Local LLM on iPad Pro: Apple Silicon Guide
Run local LLM on iPad Pro with Apple MLX. Explore M2 vs M4 benchmarks, 120 GB/s bandwidth, 16GB RAM budgets, and 7B model execution offline.
Run SmolVLM Offline on iOS: Edge Vision Guide
Run SmolVLM and SmolVLM2 offline on iOS. Explore 4-bit Apple MLX benchmarks, SigLIP patch encoding, RAM budgets, and private on-device visual OCR.
Llama 3.2 on iPhone: Local Chat & Benchmarks
Run Llama 3.2 locally on iPhone with zero cloud latency. Compare 1B vs 3B MLX benchmarks, memory usage, token speeds, and private chat setups.
Quantization 4-Bit vs 8-Bit: Mobile LLM Guide
Compare 4-bit and 8-bit quantization for mobile LLMs on Apple Silicon. Analyze RAM limits, jetsam kills, perplexity loss, and token speeds.
Apple Silicon Memory Bandwidth for LLMs: The Hardware Guide
See how memory bandwidth governs local LLM speeds on Apple Silicon, from iPhone A18 Pro to M4, and how unified memory eliminates PCIe bottlenecks.
Best Open Source Models for iPhone: Offline LLM Guide
Compare the best open-source models for iPhone: Llama 3.2, Qwen 2.5, and SmolLM2 running offline with 4-bit MLX quantization on Apple Silicon.
Best Local Vision Models for iPhone: Offline VLM Benchmarks
Compare SmolVLM2, Qwen2-VL, and MobileVLM for on-device visual OCR and document reasoning directly on Apple Silicon without cloud APIs.
How to Run LLMs Locally on iPhone (No Cloud, 100% Offline)
Learn how to run 1.5B to 7B LLMs directly on your iPhone with Apple Silicon, MLX 4-bit quantization, and Metal shaders without internet.
Running DeepSeek-R1 on iPhone: Offline Reasoning Benchmarks
Detailed benchmarks running DeepSeek-R1 Distill on iPhone 16 Pro and M4 iPad. Memory footprint, tokens/sec, and private chain-of-thought.
Why Offline AI Chatbots on iPhone Beat Cloud Models
Understand why running local LLMs on your iPhone protects your sensitive data from training pools, server leaks, and unexpected downtime.
About Lapis
Lapis is a native iOS and iPadOS app that runs foundation models like DeepSeek-R1 and Qwen 2.5 directly on Apple Silicon, completely offline with zero telemetry.