ALL TRANSMISSIONS
SEC.12//FIELD REPORT
RUNANYWHEREY COMBINATORIOSON-DEVICE AIPRIVACYGEMMA 4

Featured by RunAnywhere (YC): Bringing 15+ On-Device AI Models to iOS in Six Weeks

LLM HUB TEAM2026-09-145 min read
VERIFIED ON-DEVICE
Featured by RunAnywhere (YC): Bringing 15+ On-Device AI Models to iOS in Six Weeks

In the rapidly expanding ecosystem of artificial intelligence, most consumer applications have become thin client wrappers around expensive, cloud-hosted API servers. While cloud LLMs offer convenience for developers, they present fundamental compromises for users: privacy exposure, recurring subscription costs, latency overhead, and complete dependency on an active internet connection.

LLM Hub was founded on an unapologetic counter-thesis: your personal intelligence belongs on your personal hardware.

Recently, Y Combinator-backed RunAnywhere published an in-depth official case study analyzing how LLM Hub ported its full suite of 15+ on-device AI models from Android to iOS in under six weeks. Here is the technical breakdown of how we achieved production-grade, zero-cloud mobile AI at scale.


The Challenge: Bringing Heavyweight LLMs to Mobile Hardware

When developers talk about local AI, they often imagine desktop computers with 24GB VRAM graphics cards running heavy quantized weights. Making models execute smoothly on consumer mobile smartphones—with constrained unified RAM, thermal throttling, and aggressive mobile OS process killers—requires an entirely different engineering mindset.

Key constraints we faced when engineering the iOS release:

  1. iOS Memory Pressure Limits: iOS enforces strict per-app memory limits. Exceeding available memory triggers immediate jetsam kills by the operating system kernel.
  2. Compute Heterogeneity: Apple Silicon utilizes unified memory shared dynamically between the CPU, Metal GPU, and Apple Neural Engine (ANE). Maximizing tokens-per-second required direct Metal kernel bindings.
  3. Multi-Modal Model Breadth: LLM Hub is not just a chat app; it powers text chat, code execution via Vibes Coder, voice transcription with Whisper ASR, and local music generation with Google Magenta Realtime 2.

Enter RunAnywhere: The Six-Week Sprint

To bridge our native Android C++ engine over to iOS swiftly, we leveraged RunAnywhere’s cross-platform on-device AI SDK. RunAnywhere’s orchestration layer provided the low-level memory paging and hardware dispatch mechanisms required to run heterogeneous open-weight models on Apple Metal without reinventing low-level shaders.

+----------------------------------------------------------------+
|                        LLM Hub iOS UI                          |
|                  (100% Pure Native Swift & SwiftUI)            |
+-------------------------------+--------------------------------+
                                |
                                v
+----------------------------------------------------------------+
|                    RunAnywhere SDK Runtime                     |
|           • Dynamic Memory Paging & KV Cache Control           |
|           • Quantized Weight Streaming (2-bit to 4-bit)        |
|           • Multi-Model State Manager                          |
+---------------+-------------------------------+----------------+
                |                               |
                v                               v
+-------------------------------+ +------------------------------+
|     Apple Metal Shaders       | |     Apple Neural Engine      |
|    (LLM Inference, Gemma 4,   | |   (Whisper ASR,              |
|    Granite 4.2, LFM 2.5)      | |   Magenta Realtime 2 Audio)  |
+-------------------------------+ +------------------------------+

Key Technical Milestones Achieved:

  • Quantized Memory Efficiency: Utilizing modern quantization formats (AWQ, GGUF, and 4-bit integer weights), 3B to 4B parameter models run smoothly within a 2.2GB memory footprint on iPhone 15 Pro, iPhone 16, and newer iPads.
  • Cold-Start Optimization: Model weights load in under 1.8 seconds using memory-mapped I/O (mmap), allowing users to switch effortlessly between models like Gemma-4 and IBM Granite 4.2.
  • Zero Background Telemetry: Unlike other mobile apps that silently phone home to third-party tracking endpoints, LLM Hub contains zero tracking SDKs and zero telemetry.

Supported State-of-the-Art On-Device AI Models

To deliver uncompromising on-device performance and complete local privacy, LLM Hub focuses on top-tier, transparent, open-weight architectures engineered by world-class research labs:

Model Architecture Developer / Lab Parameter Size Primary Use Case on LLM Hub
Gemma-4 Google DeepMind e2/4b, 12b, 26be4b, 31b General reasoning, deep logical synthesis & multilingual chat
IBM Granite 4.2 IBM Research 3B - 8B Enterprise workflows, structured extraction & coding
LiquidAI LFM 2.5 Liquid AI 1.3B - 3B Ultra-low power, adaptive continuous-sequence modeling
Muse Glimmer Meta 30B Creative prose, summarization & roleplay personas
GPT OSS 20B Open Weights 20B (MoE Quant) Complex problem solving & programmatic analysis
Ministral 3 Mistral AI 3B - 8B High-speed mobile assistant with strict European safety standards
Magenta Realtime 2 Google Research Audio Engine Real-time on-device offline music & beat synthesis
Whisper ASR OpenAI Tiny / Base Offline voice-to-text with 99% accuracy

Why True Offline Privacy Matters

When you type questions or upload documents into cloud AI services, your proprietary data is transmitted over public networks, stored in centralized data centers, and frequently indexed to train future commercial models.

With LLM Hub:

  • Airplane Mode by Design: Disconnect your Wi-Fi, turn on Airplane Mode, and the entire app continues to operate flawlessly.
  • Confidential Document Analysis: Review financial statements, medical summaries, and private source code without legal exposure.
  • No Account Required: You never need an email, phone number, or credit card to access intelligence.

Read the Full Case Study on RunAnywhere

We are proud to have our engineering journey showcased by the RunAnywhere engineering team. To read the complete technical case study, including compiler optimizations and benchmarks on Apple A17/A18 Pro chips, visit the official link:

👉 Read the Official RunAnywhere Case Study

Download LLM Hub today on iOS and Android to experience the future of autonomous, sovereign, on-device AI.

Frequently Asked Questions

Q.01

What is RunAnywhere and why did LLM Hub partner with them?

RunAnywhere is a Y Combinator-backed infrastructure platform specializing in cross-platform on-device AI runtime acceleration. Partnering with RunAnywhere allowed LLM Hub to bring its multi-model offline architecture from Android to iOS in just six weeks.

Q.02

Which AI models run offline on iOS through LLM Hub?

LLM Hub runs Gemma-4 (e2/4b, 12b, 26be4b, 31b), IBM Granite 4.2, LiquidAI LFM 2.5, Muse Glimmer, GPT OSS 20B, Ministral 3, Google Magenta Realtime 2 for audio/music generation, and Whisper for local speech recognition.

Q.03

Does LLM Hub send any data or prompts to external servers?

Zero. Every single calculation, token generation, and audio waveform synthesis executes directly on your iPhone's Apple Silicon Metal GPU and Neural Engine. There is no cloud fallback, no tracking, and no external telemetry.

ZERO CLOUD // ZERO TRACKING

Experience True Private AI on Your Device

Run Gemma-4, IBM Granite 4.2, LiquidAI, Whisper, and offline image/music generation right from your pocket. Available on Android and iOS.