ALL TRANSMISSIONS
SEC.12//FIELD REPORT
ON-DEVICE AIQUANTIZATIONLLMSEXPLAINERGUIDE

Quantization Explained: How 7B Models Fit on a Phone

LLM HUB TEAM2026-10-016 min read
VERIFIED ON-DEVICE
Quantization Explained: How 7B Models Fit on a Phone

Seven billion parameters sounds like something that belongs in a data center, not in your pocket. And at full precision, it would be: a 7-billion-parameter language model stored the normal way needs about 14 gigabytes of memory, which is more RAM than most phones have in total. Yet thousands of people run models exactly this size on their phones every day, fully offline, answering questions and drafting emails with no internet in sight.

The trick that makes this possible is called quantization. It is the single most important technique in on-device AI, the reason a modern phone can hold a capable AI model, and it is worth understanding if you run — or want to run — AI locally on your phone.

The memory problem, in plain terms

A language model is a huge collection of numbers called weights. Each weight is a parameter, and together they encode everything the model "knows." When you chat with a model, your phone loads all those weights into memory and does math on them to predict the next word.

Traditionally, every weight is stored as a 16-bit floating-point number — 2 bytes. The math is brutal:

  • 1 billion parameters × 2 bytes = 2 GB
  • 7 billion parameters × 2 bytes = 14 GB
  • 70 billion parameters × 2 bytes = 140 GB

A flagship phone in 2026 might have 12 or 16 GB of RAM total, shared with the operating system and every running app. A 14 GB model simply does not fit. Something had to give.

What quantization actually does

The key insight: models do not need all that precision. Most of a weight's 16 bits of precision are wasted — the model's behavior barely changes if you round its numbers.

Quantization rounds every weight to a lower precision, like converting measurements from millimeters to whole centimeters. The common formats:

Format Bytes per weight 7B model size
FP16 (original) 2 ~14 GB
INT8 1 ~7 GB
INT4 0.5 ~3.5 GB

Going from 16-bit floats to 4-bit integers shrinks the model to roughly a quarter of its size. In practice a quantized 7B model file lands around 4-5 GB including metadata and working overhead, which fits comfortably on most modern phones and leaves room for the operating system to breathe.

The 8-bit (INT8) format sits in the middle: about 7 GB for a 7B model. It is still too big for many phones, which is why 4-bit has become the standard for mobile. The quality difference between INT8 and INT4 is real but modest — and for the kinds of things people do on phones, rarely noticeable.

How quality survives rounding

If you naively rounded every weight to 4 bits, quality would collapse. A 4-bit number can only hold 16 distinct values, while a 16-bit float can hold over 65,000. The clever part of quantization is how it decides which 16 values each weight gets.

Modern quantization works in small groups. Instead of one rounding scheme for the entire model, weights are divided into blocks (typically 32 at a time), and each block gets its own scale and range. Most weights in a block cluster around zero, so the scheme spends its limited values on the fine detail near zero and lets the rare large values sit a bit coarser.

The result: a 4-bit quantized model usually scores within a few percentage points of its full-precision parent. In everyday use — chatting, summarizing articles, drafting messages, translating — the difference is genuinely hard to spot. This is not a hacky approximation; it is a principled compression, and it is why the entire on-device AI ecosystem runs on quantized models.

One caveat is worth knowing: quantization stacks multiplicatively with small models. A 70B model quantized to 4 bits stays brilliant. A 1B model quantized to 4 bits is already small and loses more of its edge. If a tiny model feels dumbed-down in an app, its size is usually the bigger factor, not the quantization — and stepping up one model size helps more than hunting for a higher-precision version.

Q4, Q5, Q8: reading the labels

When you download a model for on-device use, you will see labels like Q4_0, Q5_K_M, or Q8_0. These are quantization levels:

  • Q4 (4-bit): the standard choice for phones. Smallest files, fastest loading, quality within a few percent of full precision.
  • Q5 (5-bit): a middle ground. Slightly larger, slightly better quality. Worth it on phones with 12 GB+ RAM if you want the best answers from a given model.
  • Q8 (8-bit): near-full-precision quality at half the size. Mostly useful on high-RAM devices or desktops; usually overkill for a phone.

The middle letters and numbers (the _K_M suffixes) refer to the exact quantization method — different schemes for different layer types. You can ignore them. Just pick Q4 if you want the balanced default, and Q5 if your phone has plenty of RAM and you want the best quality available.

Why this matters for your phone

Quantization is what turned on-device AI from a research demo into something you can use daily. The consequences are practical:

Model choice becomes storage choice. Because 1 billion parameters cost about half a gigabyte at Q4, you can estimate any model: a 3B model is ~2 GB, a 7-8B model is ~4-5 GB. That simple rule lets you pick the best model that fits your phone's free space.

Phones that couldn't, now can. Every generation of phone hardware makes quantized models faster, but quantization itself is what made them possible at all. Without it, on-device AI would still be limited to tiny models that can barely hold a conversation.

Everything stays private. This is the payoff that matters most. Because the model is small enough to live on your phone, your prompts never leave the device. No servers, no logs, no accounts needed. Quantization is a compression story, but its real achievement is privacy: the reason a capable AI assistant can be completely offline is that it first got small enough to fit.

Picking a quantized model: the quick version

If all of this is new, here is the decision in one paragraph. On a phone with 8 GB of RAM, download a 3-4B model at Q4 and you will have a fast, capable offline assistant. On a phone with 12 GB or more, try a 7-8B model at Q4 for noticeably smarter answers. Short on storage? Drop one model size rather than hunting exotic formats — the size difference dwarfs the precision difference.

Apps like LLM Hub handle all of this behind the scenes: models are already quantized and tuned for phones, so downloading one is a single tap. But now you know what those Q4 labels mean, why a 7-billion-parameter model fits in your pocket, and why the entire industry rounds its AI to fit in your hand.

Frequently Asked Questions

Q.01

What is quantization in AI models?

Quantization is a compression technique that reduces the precision of a model's numbers (weights), typically from 16-bit floats to 4-bit integers. It cuts a model's size to about a quarter with only a small drop in quality — it is the main reason phones can run capable 7B language models offline.

Q.02

How big is a quantized 7B model?

At 4-bit quantization, a 7-billion-parameter model needs roughly 3.5-5GB of storage and a similar amount of RAM while running. At full 16-bit precision it would need about 14GB, which is more than most phones can hold — that is why quantization is essential.

Q.03

Does quantization make AI models worse?

Slightly. A 4-bit quantized model typically scores within a few percent of its full-precision version on common benchmarks. For everyday tasks like chatting, summarizing, and translating, the difference is usually imperceptible — though very small models quantized aggressively can feel noticeably less capable.

Q.04

What do Q4, Q8 and Q5 mean in model downloads?

They are quantization levels: Q4 means 4-bit weights, Q8 means 8-bit, and Q5 is in between. Lower numbers mean smaller files and faster loading but slightly lower quality. Q4 is the standard choice for phones; Q5 or Q8 make sense on phones with lots of RAM if you want maximum quality.

Q.05

Can I quantize a model myself for my phone?

In theory yes — tools like llama.cpp can quantize models — but in practice you don't need to. Apps like LLM Hub ship pre-quantized models that are already tuned for phones, so you just download and chat.

ZERO CLOUD // ZERO TRACKING

Experience True Private AI on Your Device

Run Gemma-4, IBM Granite 4.2, LiquidAI, Whisper, and offline image/music generation right from your pocket. Available on Android and iOS.