You recorded a 40-minute interview on your phone. Now you're on a train with no signal, and the deadline is tonight. If that audio has to go to the cloud first, you're stuck. If your phone can transcribe it offline, you spend the ride editing a draft instead of staring out the window.
This is exactly what on-device speech transcription does. Models like Whisper, the open-source speech recognition model, now run directly on phones. Record a voice memo, a lecture, an interview, or a meeting, and your phone turns it into text with no internet connection, no subscription, and no audio ever leaving your device.
This guide explains how offline transcription works, which model sizes make sense on a phone, and how to get transcripts worth reading.
How offline transcription works
Speech-to-text is a neural network trained to listen to audio and output words. When you tap "transcribe" in an offline app, here's what happens on your phone:
- Audio preprocessing. The app converts your recording into a standard format: a spectrogram, which is a visual representation of the sound frequencies over time. Think of it as the model's eyes for sound.
- The neural network runs. Whisper's encoder listens to the spectrogram and its decoder predicts the spoken text, token by token. This is where the phone's chip does the heavy lifting, often with help from the NPU.
- Text output. You get a transcript, usually with timestamps per segment, which you can edit, copy, or export.
In a cloud app, steps 2 and 3 happen on a remote server after your audio uploads. Offline, everything happens in your pocket. A 10-minute voice memo might take anywhere from one minute to ten minutes to transcribe depending on the model size and your phone.
The models themselves are the same architecture that powers popular cloud transcription services, compressed to fit on a phone. OpenAI released Whisper as open source, which is why it shows up in so many offline apps. It's genuinely good technology, not a cut-down demo.
Choosing a model size: the real tradeoff
Whisper comes in several sizes, and the name tells you roughly what to expect:
- Tiny (~39M parameters, ~75 MB). Fastest, smallest. Good enough for rough drafts and keyword searching, but it makes noticeably more errors, especially with accents and noisy audio.
- Base (~74M parameters, ~145 MB). A step up. Decent for clear voice memos and personal notes.
- Small (~244M parameters, ~465 MB). The phone sweet spot for most people. Strong accuracy on clear speech, manageable speed on modern phones.
- Medium (~769M parameters, ~1.5 GB). Best accuracy you'd reasonably run on a flagship phone. Slow on older devices, fine on recent flagships with a good NPU.
- Large (~1.5B+ parameters). Desktop and server territory. You can technically try it on a phone, but it's slow enough that you won't want to.
Here's the practical advice: start with small. If your recordings are clear and your phone is recent, small gives you transcripts that need only light cleanup. If you're transcribing hours of audio regularly, tiny or base in "rough draft" mode saves enormous time, and you can rerun the important parts with a bigger model later.
Storage matters too. A small model plus the app might cost you half a gigabyte. That's one podcast episode. Medium pushes past a gigabyte, so make sure you have room.
What it handles well, and what it doesn't
Let's be honest about the limits, because they're mostly about your recording, not the model.
What it handles well:
- Clear speech in a quiet room: lectures, interviews, voice memos, dictated notes.
- Many languages: Whisper handles dozens, with automatic language detection.
- Punctuation and capitalization: Whisper adds both, so transcripts are readable, not raw word streams.
- Timestamps: segments are time-stamped, so you can jump back to the audio at any point.
What trips it up:
- Background noise. Cafes, traffic, wind. The model can only work with what the microphone captured. A quiet room is worth more than a bigger model.
- Overlapping speakers. Two people talking at once confuses every speech model, cloud or on-device. Speaker diarization (who said what) is a separate, harder problem; most phone apps label speakers imperfectly or not at all.
- Heavy accents and mumbling. Models are trained on lots of data, but far-from-training-distribution speech still degrades accuracy.
- Jargon and names. Uncommon technical terms, product names, and proper nouns are the most common transcription errors. Keep an eye on them when editing.
The single biggest quality lever is free: record well. Hold the phone reasonably close, pick the quiet room, use airplane mode so notifications don't interrupt. A good recording on the tiny model beats a bad recording on the medium model.
Practical workflows people actually use
Journalists and interviewers. Record the interview, transcribe on the device, pull quotes straight from the text. The privacy angle matters here: a source who spoke in confidence shouldn't have their voice uploaded to a cloud server. On-device transcription keeps the audio exactly where it was recorded.
Students. Record lectures (where permitted), transcribe them, then ask an on-device AI assistant to summarize the transcript into study notes. The whole pipeline runs offline: transcribe, summarize, quiz yourself. Some students run this entire loop in airplane mode on the commute.
Voice memo people. If you think out loud into your phone instead of typing notes, transcription turns those memos into searchable text. Search "dentist" instead of scrubbing through audio. Pair it with a local notes app and your voice memos become a personal knowledge base.
Meetings and calls. Record with consent, transcribe after, and get action items. On-device means the meeting content never touches a third-party server, which matters for sensitive business or client work.
Accessibility. Deaf and hard-of-hearing users get live-ish captions for recorded audio. Real-time live captioning needs streaming models, but transcribing a recording right after is already transformative.
Privacy: the part cloud apps can't match
Cloud transcription services are convenient, and many are fine products. But the architecture is what it is: your audio uploads to someone else's server, gets transcribed there, and is stored according to that company's policies. For casual dictation, that's a reasonable tradeoff. For the following, it isn't:
- Confidential interviews and source material
- Medical, legal, or therapy-related recordings
- Internal business meetings and strategy discussions
- Anything involving children or vulnerable people
With on-device transcription, the audio never leaves the phone. The transcript is as private as the recording. There is no retention policy to read, no subpoena risk against a third party, no "improve our models" clause quietly training on your voice. Apps like LLM Hub go further: no accounts, no tracking, no analytics, so there's no metadata trail either. For genuinely sensitive audio, offline isn't just convenient, it's the only setup that keeps your data yours.
How to do it in LLM Hub
LLM Hub includes an on-device Whisper transcriber as one of its built-in tools, alongside chat, image generation, translation, and the rest. The workflow is simple:
- Open the transcriber and pick your Whisper model size (small is the default recommendation).
- Record directly in the app, or import an existing audio file or voice memo.
- The app transcribes entirely on-device. No internet needed at any point.
- Edit the transcript, copy it, or export it to your notes app.
Because LLM Hub keeps everything local with no accounts and no tracking, the audio and the transcript live on your phone and nowhere else. It works the same in airplane mode as it does on Wi-Fi.
Getting clean transcripts: a quick checklist
Before you record, run through this list. It takes ten seconds and it matters more than which model you pick:
- Quiet location. Background noise is the number one transcript killer.
- Phone close to the speaker. Two to three feet is ideal. Across the room is not.
- Airplane mode. Kills notifications, saves battery, and guarantees the app stays fully offline.
- Set the language manually if you know it. Skipping auto-detection speeds things up.
- Pick small for everyday use, medium for important recordings on a recent phone.
- Proof the names. Jargon and proper nouns are where errors cluster. A quick skim of the transcript catches most of them.
- Keep the audio. Transcripts are drafts, not records. Keep the original recording until the transcript is verified.
The bottom line
Offline Whisper transcription on your phone is no longer a tech demo. It's a daily-use tool: fast enough, accurate enough on clean audio, and private by architecture rather than by promise. The cloud version is still better for the hardest audio and for speaker separation, but for interviews, lectures, voice memos, and meetings recorded in reasonable conditions, your phone can do the job alone, in airplane mode, with nothing leaving the device.
If your audio is sensitive, offline isn't just the convenient option. It's the right one.

