You downloaded an offline AI app, tapped "download model," and watched a 4GB file crawl onto your phone. Or maybe you are still deciding whether your phone can handle it at all. The two numbers that matter are RAM and storage, and unlike cloud AI — where some data center worries about hardware — on-device AI puts the hardware question squarely on your phone.
The good news: the math is simple, and most modern phones clear the bar. This guide walks through exactly how much memory and storage you need, why quantization changed everything, and how to pick the right model for the phone in your pocket.
The one formula that decides everything
Every language model is a giant collection of numbers called parameters (weights). To generate text, your phone loads those numbers into RAM and does math on them. So the model's memory footprint is essentially:
Memory needed ≈ number of parameters × bytes per parameter + working memory
The two variables you control are the model's size (parameter count) and its precision (how many bytes each parameter takes). Here is what precision means in practice for a 1-billion-parameter model:
- FP16 (16-bit): 2 bytes per parameter → about 2GB
- INT8 (8-bit): 1 byte per parameter → about 1GB
- INT4 (4-bit): half a byte per parameter → about 0.5GB
INT4 is the format most phone apps use, thanks to a technique called quantization (covered in detail in our companion article on how 7B models fit on a phone). It shrinks models to roughly a quarter of their original size with only a small quality tradeoff — and it is the entire reason phones can run serious language models at all.
What different model sizes actually need
Here is a practical table of what real on-device models demand, assuming 4-bit quantization, which is the standard for mobile apps:
| Model size | Model file on disk | RAM while running | Comfortable on |
|---|---|---|---|
| 1B parameters | ~0.5–0.7 GB | ~1–1.5 GB | 4GB+ RAM phones |
| 3B parameters | ~1.5–2 GB | ~2.5–3.5 GB | 6–8GB RAM phones |
| 7–8B parameters | ~4–5 GB | ~5–6 GB | 12GB+ RAM phones |
| 12B+ parameters | ~6–8 GB+ | ~8–10 GB+ | 16GB RAM flagships only |
Two things to notice. First, the "RAM while running" column is bigger than the file size, because the model needs working memory on top of its weights: the KV cache (a scratchpad of the conversation so far, which grows with context length), the runtime's own overhead, and — critically — the memory your operating system and other apps need to stay alive. A good rule of thumb: leave 1.5–2GB of headroom above the model's footprint.
Second, these are rough numbers, not exact specs. Different runtimes and model architectures have different overheads. Treat the table as a map, not a contract.
Where most phones land in 2026
Phone RAM has grown steadily, and the mid-range sweet spot now sits at 8GB. Here is the honest breakdown by tier:
4–6GB RAM (older or budget phones): Stick to 1B-class models. A model like SmolLM2-1.7B or a 1B Gemma-class model at 4-bit quantization fits comfortably. Close background apps before a long session, keep conversations short, and you will get genuinely useful results — these small models are far more capable than their size suggests.
8GB RAM (the mainstream): This is the comfortable home of 3B models. A 3B Llama-class or Phi-class model at 4-bit quantization runs well with room to spare for the OS. You can also try a 7–8B model, but expect the OS to start evicting background apps aggressively, and longer conversations may slow down as the KV cache grows.
12GB+ RAM (flagships and gaming phones): 7–8B models are realistic here. This is where offline AI gets genuinely impressive: an 8B model at 4-bit is the largest class of model that runs practically on any current phone, and on 12GB+ devices it has enough breathing room for longer contexts without the phone struggling.
16GB RAM (ultra flagships): The playground for 12B-class experiments. These exist, and they run, but honestly an 8B model on 12GB gives a better experience per watt than a 12B model squeezed onto a phone that spends half its effort managing memory.
Storage: the forgotten half of the question
RAM decides which models can run. Storage decides which models you can keep. Each downloaded model is a file that lives on your phone permanently until you delete it — and unlike streaming a movie, you need the whole file.
At 4-bit quantization, budget roughly 0.5GB of storage per billion parameters: a 3B model is ~1.5–2GB, an 8B model is ~4–5GB. That does not sound like much on a 128GB phone, until you remember your photos, apps, and offline maps are competing for the same space.
Practical advice:
- Download on Wi-Fi. Model files are big. Your mobile data plan will not thank you for a 5GB download.
- Keep one or two models, not ten. A 3B general model plus a specialist (say, a coding model) covers nearly everything. Ten downloaded models you never open are just storage debt.
- Keep 5–10GB free. Phones slow down and updates fail when storage runs critically low. Treat model files like any other large media.
- Delete ruthlessly. Finished experimenting with a model? Delete the file. Re-downloading later costs minutes on Wi-Fi, not money.
Why RAM is the gatekeeper, not the chip
People obsess over which chip their phone has — and chips do matter for speed. But here is the hierarchy of what actually decides your experience:
- RAM determines which models can run at all. If the model does not fit, nothing else matters.
- Memory bandwidth (how fast data moves between RAM and processor) determines generation speed more than raw compute does. Running a model is mostly about shuttling billions of numbers around, not crunching them.
- NPU/GPU capability determines how efficiently that shuttling turns into tokens per second.
- Thermals determine how long peak speed lasts before the phone throttles.
This is why a phone with 12GB of RAM and a mid-range chip often gives a better offline-AI experience than a flagship chip paired with 8GB of RAM: the first phone can load a bigger model, and model size affects answer quality far more than a 20% speed difference does.
The context-length tax
One more RAM consumer that surprises people: conversation length. As you chat, the model stores a representation of everything said so far — the KV cache. A short exchange costs almost nothing. But paste a long document into the chat, or carry on a very long conversation, and the cache can grow to hundreds of megabytes or more.
If a long session starts to slow down, that is usually why. Starting a fresh chat instantly frees that memory. Some apps also let you set a maximum context length — shorter limits mean less RAM pressure and snappier responses, at the cost of the model "forgetting" earlier messages.
Battery: the question everyone asks next
Running a model keeps your phone's processor busy, so yes — an active offline-AI session drains the battery faster than reading the news, roughly in the league of mobile gaming. A few honest notes:
- Bigger models drain faster. A 1B model sips; an 8B model gulps. Match the model to the task.
- Idle cost is near zero. Unlike cloud apps, there is no radio constantly syncing. When you are not generating text, the app barely uses power.
- Shorter answers save battery. Every token costs energy. Concise prompts and concise answers are the greenest way to use on-device AI.
Choosing your model: a simple decision tree
Forget the spec-sheet anxiety. Answer three questions:
- How much RAM does your phone have? Check Settings → About. Under 8GB: 1–3B models. 8–12GB: 3B comfortably, 7–8B experimentally. 12GB+: 7–8B comfortably.
- What do you actually need it for? Quick answers, translation, and drafting: 1–3B is plenty. Coding help, complex reasoning, long documents: go as big as your RAM allows.
- How much free storage do you have? If you are under 10GB free, pick one small model and delete it when done.
That is genuinely the whole decision. The offline-AI world rewards pragmatism: the best model is the one that fits your phone, answers your question, and leaves your battery alive for the rest of the day.
How LLM Hub handles this
LLM Hub ships with 15+ downloadable models across these size classes, so you can match the model to your phone instead of guessing. The app downloads model files to your device over Wi-Fi, runs them fully offline with no account and no tracking, and you can delete any model file at any time to reclaim storage. Start with a 3B-class model if your phone has 8GB of RAM — it is the sweet spot of quality, speed, and battery life for most people — and only step up to larger models if you have the RAM and storage to spare.
Your phone is already a more capable AI computer than most people realize. The trick is simply knowing how much headroom you have — and now you do.

