In the rapidly expanding ecosystem of artificial intelligence, most consumer applications have become thin client wrappers around expensive, cloud-hosted API servers. While cloud LLMs offer convenience for developers, they present fundamental compromises for users: privacy exposure, recurring subscription costs, latency overhead, and complete dependency on an active internet connection.
LLM Hub was founded on an unapologetic counter-thesis: your personal intelligence belongs on your personal hardware.
Recently, Y Combinator-backed RunAnywhere published an in-depth official case study analyzing how LLM Hub ported its full suite of 15+ on-device AI models from Android to iOS in under six weeks. Here is the technical breakdown of how we achieved production-grade, zero-cloud mobile AI at scale.
The Challenge: Bringing Heavyweight LLMs to Mobile Hardware
When developers talk about local AI, they often imagine desktop computers with 24GB VRAM graphics cards running heavy quantized weights. Making models execute smoothly on consumer mobile smartphones—with constrained unified RAM, thermal throttling, and aggressive mobile OS process killers—requires an entirely different engineering mindset.
Key constraints we faced when engineering the iOS release:
- iOS Memory Pressure Limits: iOS enforces strict per-app memory limits. Exceeding available memory triggers immediate jetsam kills by the operating system kernel.
- Compute Heterogeneity: Apple Silicon utilizes unified memory shared dynamically between the CPU, Metal GPU, and Apple Neural Engine (ANE). Maximizing tokens-per-second required direct Metal kernel bindings.
- Multi-Modal Model Breadth: LLM Hub is not just a chat app; it powers text chat, code execution via Vibes Coder, voice transcription with Whisper ASR, and local music generation with Google Magenta Realtime 2.
Enter RunAnywhere: The Six-Week Sprint
To bridge our native Android C++ engine over to iOS swiftly, we leveraged RunAnywhere’s cross-platform on-device AI SDK. RunAnywhere’s orchestration layer provided the low-level memory paging and hardware dispatch mechanisms required to run heterogeneous open-weight models on Apple Metal without reinventing low-level shaders.
+----------------------------------------------------------------+
| LLM Hub iOS UI |
| (100% Pure Native Swift & SwiftUI) |
+-------------------------------+--------------------------------+
|
v
+----------------------------------------------------------------+
| RunAnywhere SDK Runtime |
| • Dynamic Memory Paging & KV Cache Control |
| • Quantized Weight Streaming (2-bit to 4-bit) |
| • Multi-Model State Manager |
+---------------+-------------------------------+----------------+
| |
v v
+-------------------------------+ +------------------------------+
| Apple Metal Shaders | | Apple Neural Engine |
| (LLM Inference, Gemma 4, | | (Whisper ASR, |
| Granite 4.2, LFM 2.5) | | Magenta Realtime 2 Audio) |
+-------------------------------+ +------------------------------+
Key Technical Milestones Achieved:
- Quantized Memory Efficiency: Utilizing modern quantization formats (AWQ, GGUF, and 4-bit integer weights), 3B to 4B parameter models run smoothly within a 2.2GB memory footprint on iPhone 15 Pro, iPhone 16, and newer iPads.
- Cold-Start Optimization: Model weights load in under 1.8 seconds using memory-mapped I/O (
mmap), allowing users to switch effortlessly between models like Gemma-4 and IBM Granite 4.2. - Zero Background Telemetry: Unlike other mobile apps that silently phone home to third-party tracking endpoints, LLM Hub contains zero tracking SDKs and zero telemetry.
Supported State-of-the-Art On-Device AI Models
To deliver uncompromising on-device performance and complete local privacy, LLM Hub focuses on top-tier, transparent, open-weight architectures engineered by world-class research labs:
| Model Architecture | Developer / Lab | Parameter Size | Primary Use Case on LLM Hub |
|---|---|---|---|
| Gemma-4 | Google DeepMind | e2/4b, 12b, 26be4b, 31b | General reasoning, deep logical synthesis & multilingual chat |
| IBM Granite 4.2 | IBM Research | 3B - 8B | Enterprise workflows, structured extraction & coding |
| LiquidAI LFM 2.5 | Liquid AI | 1.3B - 3B | Ultra-low power, adaptive continuous-sequence modeling |
| Muse Glimmer | Meta | 30B | Creative prose, summarization & roleplay personas |
| GPT OSS 20B | Open Weights | 20B (MoE Quant) | Complex problem solving & programmatic analysis |
| Ministral 3 | Mistral AI | 3B - 8B | High-speed mobile assistant with strict European safety standards |
| Magenta Realtime 2 | Google Research | Audio Engine | Real-time on-device offline music & beat synthesis |
| Whisper ASR | OpenAI | Tiny / Base | Offline voice-to-text with 99% accuracy |
Why True Offline Privacy Matters
When you type questions or upload documents into cloud AI services, your proprietary data is transmitted over public networks, stored in centralized data centers, and frequently indexed to train future commercial models.
With LLM Hub:
- Airplane Mode by Design: Disconnect your Wi-Fi, turn on Airplane Mode, and the entire app continues to operate flawlessly.
- Confidential Document Analysis: Review financial statements, medical summaries, and private source code without legal exposure.
- No Account Required: You never need an email, phone number, or credit card to access intelligence.
Read the Full Case Study on RunAnywhere
We are proud to have our engineering journey showcased by the RunAnywhere engineering team. To read the complete technical case study, including compiler optimizations and benchmarks on Apple A17/A18 Pro chips, visit the official link:
👉 Read the Official RunAnywhere Case Study
Download LLM Hub today on iOS and Android to experience the future of autonomous, sovereign, on-device AI.

