Short answer: It depends entirely on the app. Voice to text always records audio briefly to convert it, but whether that audio is kept afterward varies. Some services retain recordings and transcripts to improve their systems; others delete audio immediately. Check the retention section of the privacy policy, not the marketing page.
It is a reasonable question to get nervous about. You are holding down a key and talking to your computer about a client, a diagnosis, a salary negotiation or a chapter of your novel. Something is listening. The honest answer to "is it storing that" is that the answer is different for every app, and most people never find out which kind they are using.
This guide explains what actually happens to your voice between the moment you speak and the moment text appears, what the meaningfully different privacy models are, how to read a policy and get a real answer in about ninety seconds, and what you should simply never dictate no matter which app you trust.
What happens between your mouth and the text
Every voice-to-text system, without exception, follows the same four steps.
- Capture. Your microphone converts sound into audio data. This happens on your device. There is no way to skip it, because there is nothing to transcribe otherwise.
- Buffer. That audio sits somewhere for a moment. In memory, or briefly in a temporary file, while the system decides what to do with it.
- Recognize. The audio is turned into words, either by something running on your device or by something running on a server.
- Insert. The resulting text is placed where your cursor is, and the audio is no longer needed for the task.
The privacy question is not really about steps one and two. All software that transcribes your voice records your voice; that is what transcription is. The question is what happens at step three and, more importantly, what happens after step four. Does the audio get deleted? Does the transcript get kept? Does anything leave your device, and if so, does it stay gone?
The privacy models that actually differ
Fully on-device
Recognition happens on your machine and nothing goes anywhere. This is the strongest privacy posture available, and it is what people usually picture when they say they want "offline dictation." The tradeoffs are real: on-device recognition is generally less accurate with unusual vocabulary, punctuation and accents, and it depends on your hardware. We wrote about the practical side of this in our look at offline voice to text on Mac.
Cloud processing, nothing retained
Audio is sent to a server, converted, and the text comes back. The audio is not kept afterward, and neither is the text. You get the accuracy of a large modern transcription system without leaving a growing archive of your speech in someone's storage bucket. What makes this model trustworthy or not is whether the retention claim is written down and specific.
Cloud processing, retained for improvement
Audio, transcripts or both are stored, often with the stated purpose of improving accuracy. Sometimes there is a toggle to opt out and sometimes there is not. Sometimes human review of samples is part of it, which is usually disclosed somewhere in the policy but rarely on the homepage. This is not inherently sinister, but it is a materially different deal, and you should know if you are in it.
Deliberately archival
Meeting transcription tools are a category of their own, because storing the recording is the product. If a tool exists to give you a searchable archive of your calls, then yes, obviously, it is storing audio. That is what you bought. The thing worth checking there is who else can reach the archive and how long it lives.
How to find the real answer in ninety seconds
Skip the homepage. Marketing pages say "private" and "secure" for free, and those words are not commitments. Open the privacy policy and use your browser's find function on these terms:
- "retain" / "retention" — the core question. How long is anything kept?
- "audio" / "voice recordings" — is audio named separately from other data? Vagueness here is a signal.
- "improve" / "train" — the usual phrasing for keeping your content to make the system better.
- "human review" / "manual review" — whether people may listen to or read samples.
- "third parties" / "subprocessors" — who else touches the data on its way through.
- "delete" / "your rights" — whether you can actually get anything removed.
Two questions cut through most of the fog. First: is there a specific retention period, or just words like "as long as necessary"? Specificity is a commitment; vagueness is optionality. Second: is audio addressed separately from account data? A policy that talks at length about your email address and never once says the word "audio" has not answered your question.
One more thing worth checking on mobile, because it confuses people constantly: a third-party iPhone keyboard has to be granted something Apple calls Full Access before it can talk to the network at all. iOS shows a fairly alarming warning when you enable it. That warning is generic, applied to every keyboard equally, and says nothing about what any particular keyboard does. The permission is required for any keyboard that transcribes in the cloud. What matters is the policy behind it, not the dialog.
What is stored on your own device is a separate question
People ask about servers and forget the machine in front of them. Most dictation tools keep something locally: a history of recent transcriptions, usage statistics, a personal dictionary. That local record is usually a feature, and often a useful one, but it is worth knowing it exists, especially on a shared or work-issued computer.
Ask three things about local storage. Where does it live? Can you clear it? Is it included in your backups, and if so, where do those backups go? A transcript history synced to a cloud backup has quietly become cloud storage even if the app never sent it anywhere.
Voice Keyboard Pro keeps your history and your settings in a folder on your own Mac. That is the copy that exists, and you can delete it.
Where Voice Keyboard Pro stands
Plainly: our server stores only operational pings. No audio and no transcript content is retained on our server. The things that pass through are the operational facts needed to run the service, and the words you dictate are not among them. Your text stays on your device.
We think this is the right default rather than a premium feature, and it is why we made it the standing behavior rather than a toggle you have to find. If you want the longer version of our reasoning and setup guidance, see our guide to private voice to text on Mac.
The threat model most people actually have
It is worth being honest about what you are protecting against, because the answer changes what you should do.
If your concern is a vendor building a corpus of your speech, retention policy is the whole ballgame. Find the retention section, get a specific answer, and choose accordingly.
If your concern is a legal or contractual obligation — client confidentiality, patient information, an NDA, a regulated industry — then policy alone is not sufficient. You need the agreement your organization requires, and that is a conversation with your compliance or IT team, not something an app's marketing page can settle for you. Our post on dictation and clinical notes covers why the question is organizational rather than technical.
If your concern is people physically near you, no privacy policy helps at all. Dictating a performance review in an open-plan office is a disclosure regardless of what the server does. This is the risk people underrate the most, and the only fix is where and when you speak.
If your concern is your employer, understand that on a managed device the dictation app may be the least of it. Managed machines can have monitoring at layers well below any individual application.
Things not to dictate, regardless
Some content should not go through voice input at all, and the reason is not distrust of any particular vendor. It is that these items are both high-consequence and badly suited to speech in the first place.
- Passwords and passphrases. Never. They also transcribe terribly, so you get the worst of both.
- API keys, tokens and secrets. Long random strings are exactly what transcription is worst at, and exactly what you least want floating around.
- Full card numbers, bank account and routing numbers. Type them.
- Government identification numbers. Same reasoning.
- Anything you are contractually barred from processing with an outside service. If a contract names the systems you may use, that list is the answer.
Notice that most of that list is strings rather than sentences, and strings are the category speech handles worst anyway. If you want the mechanics of why, and the technique for the borderline cases like email addresses and URLs, we covered it in dictating email addresses and URLs. The short version is that a transcription engine is built for language, and a password is the opposite of language.
Practical hygiene that takes five minutes
- Read the retention paragraph once. Not the whole policy. The retention paragraph. Then you know.
- Check your local history. Find where it lives, decide if you want it, clear it if you do not.
- Audit your microphone permissions. On Mac, System Settings has a list of every app allowed to use the mic. Revoke the ones you do not recognize. This is worth doing annually whether or not you dictate.
- Pick a physical habit. Decide in advance which topics you will not dictate in a shared space. This costs nothing and covers the risk no policy can.
- Re-check after a major update. Policies change. A quick look after a big version bump is cheap.
Frequently asked questions
Does the built-in dictation on my phone or computer store my voice?
It depends on the vendor, the operating system version and your settings. Apple and Google both publish documentation on how their speech features handle data, and both have settings related to speech data. Check the current documentation for the version you are running rather than trusting a forum post from three years ago, because these behaviors change between releases.
Is on-device dictation always more private than cloud dictation?
For the specific question of "does audio leave my machine," yes by definition. But on-device is not automatically better on every axis: accuracy is usually lower, vocabulary handling is weaker, and a local transcript history that syncs to a cloud backup is still in a cloud. Privacy is about the whole path, not just the recognition step.
Can I dictate confidential client information?
That is your organization's call and possibly your regulator's, not ours. Ask the people who own that decision where you work. Anyone who answers "yes, definitely, for everyone" without knowing your contracts is not being careful with your career.
What does "no audio is stored" actually mean?
It should mean the audio exists only long enough to be converted, and then it is gone. What makes the claim meaningful is whether it is written in the policy in plain terms rather than implied by a tagline.
Is voice to text secure?
Security and privacy are different questions and both matter. Security is whether the data is protected in transit and at rest. Privacy is whether it is kept at all. The cleanest answer to both is data that is never retained, because nothing stored is nothing to breach.
The bottom line
Every voice-to-text app records your voice, because that is the job. The real question is whether it keeps it, and that answer lives in one paragraph of the privacy policy rather than anywhere on the marketing site. Go read that paragraph for whatever you are using today. If it is vague about audio, you have learned something.
Our position is that dictation should not cost you an archive. Voice Keyboard Pro's server stores only operational pings, with no audio and no transcript content retained. You can try it free on Mac and iPhone, with Pro at $4.99 a month or $34.99 a year, and keep the words you speak on the device you speak them into.