Short answer: Wrong-word dictation errors are rarely random. They come from homophones, unknown proper nouns, word-boundary splits, microphone problems, unclear delivery, a language mismatch, or the app's own autocorrect. Identify which one you have with a short test, then fix that specific cause.
The complaint sounds vague when you first say it out loud: dictation "gets words wrong." But look closely at a page of dictated text and the errors are not scattered noise. They cluster. The same handful of words fail repeatedly, in predictable ways, for reasons that have almost nothing to do with each other.
That is the useful insight, because it means there is no single fix. Adding a name to a personal dictionary does nothing for a bad microphone. Buying a better microphone does nothing about "their" versus "there." People try one remedy, see no improvement, and conclude that voice typing is not accurate enough for them, when in fact they applied the wrong remedy to a correctly diagnosed problem.
This guide separates the seven real causes, gives you a short test to find out which one you have, and fixes each one in the order that gives you the most improvement per minute spent.
First, why these errors are invisible
Something worth understanding before the causes: modern transcription does not misspell. It picks a real, correctly spelled word from a vocabulary, and it picks the word that best fits the sound plus the surrounding context.
The consequence is that every error it makes is a real English word sitting in a grammatical position. A spell checker will never flag it. Your eye, skimming, will often not flag it either, because your brain reads what you meant. This is different from typing errors, which usually announce themselves as red squiggles or obvious mangling.
So the first fix is not technical at all: read dictated text as a reader, not as the person who said it. Especially the negations. "Can" and "cannot" sound different enough that this is rare, but when it happens the sentence still reads perfectly and means the opposite of what you intended. That single class of error is worth more attention than all the rest combined.
The 60-second test
Open a plain text field. TextEdit on a Mac, or Notes. Do not use a browser, a chat app, or a word processor for this, because those add their own text processing and you are trying to isolate the transcription itself.
Dictate this paragraph exactly, at your normal speaking pace, in one continuous take:
They said their report would be ready by two o'clock, but I think it is going to take longer than that. If we cannot finish it today, we should tell the client tomorrow morning rather than leaving them to guess.
Now compare. What you find tells you where to go next:
- It came out perfectly. Your transcription is fine. The errors you are seeing come from proper nouns, from the app you normally dictate into, or from how you speak when you are composing rather than reading. Go to causes 2, 5 and 7.
- A few words are wrong but it is mostly right. Normal. Go to causes 1 and 3, and read the proofreading section.
- It is badly mangled, with words that are not close to what you said. This is an audio problem, not a language problem. Go straight to cause 4.
- It came out in the wrong language, or with odd accented spellings. Cause 6.
Do this test whenever accuracy suddenly gets worse. It takes a minute and it separates "something changed in my setup" from "this app is doing something to my text."
Cause 1: Homophones
Their and there and they're. To, too and two. Right, write and rite. Its and it's. Affect and effect. Peak, peek and pique. These are genuinely identical in the audio stream. No microphone upgrade will help, because there is nothing in the sound to distinguish them.
The only thing that can resolve them is context, and context is the thing most people accidentally remove.
The fix: speak in full sentences. This is counterintuitive, because when accuracy is poor the instinct is to slow down and say individual words carefully. That makes homophone errors worse, not better. A word spoken alone has no context, so the engine has nothing to work with and falls back to the most common spelling. The same word inside a complete sentence is usually resolved correctly because the surrounding grammar rules out the alternatives.
If you dictate in three-word fragments, pausing to think between each, you are giving away the one signal that fixes this class of error. Compose the sentence in your head, then say the whole thing.
The second fix: proofread for meaning, not spelling. Homophone errors will never be fully eliminated, because sometimes both readings are grammatical. Accept a small residue and catch it on the read-through.
Cause 2: Proper nouns and specialist vocabulary
This is the biggest cause of people giving up on dictation entirely, and it is the most fixable.
A transcription engine knows general language extremely well. It has no way of knowing your colleague's surname, your company's internal project names, the drug names in your specialty, the framework names in your codebase, or the town where your client's office is. Presented with an unfamiliar sound, it does the only thing it can: it substitutes the nearest common word. So a name becomes an everyday noun, every single time, in every single document.
The frustration is not that it happens once. It is that it happens identically forever, and you fix it by hand forever.
The fix: a personal dictionary. Voice Keyboard Pro's Smart Vocabulary lets you add the terms you actually use, with replacement rules, so the engine stops guessing at them. Build the list from real corrections rather than trying to anticipate it: dictate normally for two days, note the words you keep fixing, and add those. Most people find the list is shorter than expected. Twenty to thirty entries usually covers a whole professional vocabulary.
Names deserve their own attention because they also collide with autocorrect on the app side. We covered that specific fight in stopping Mac dictation from autocorrecting names. The broader mechanics of how a custom vocabulary learns your words are worth reading if you work in a jargon-heavy field.
If you fix only one thing on this page, fix this one. It is the difference between a tool that annoys you daily and one you forget you are using.
Cause 3: Word-boundary errors
Speech has no spaces in it. Where one word ends and the next begins is a decision made from context, not something present in the audio. This produces a distinctive error where the letters are roughly right but split in the wrong place.
"A nice house" and "an ice house" are the same sound. So are "some others" and "some mothers," "grade A" and "gray day," "it's a" and "it saw." The output looks strange in a way homophone errors do not, which at least makes these easy to spot.
The fix is the same as cause 1, and it is again about context. Full sentences give the engine enough surrounding grammar to split the stream correctly. Fragments do not.
There is also a delivery component. Running words together at high speed increases boundary errors, but so does the opposite. Over-enunciating, putting a hard stop between every word, actively removes the natural coarticulation the engine expects and can make output worse. The target is a clear, unhurried conversational pace. Our guide to speaking clearly for dictation goes into what that means in practice, and the short version is: talk like you are explaining something to a colleague, not like you are reading a train announcement.
Cause 4: The audio is worse than you think
If your test paragraph came out badly mangled, stop reading about language and look at the microphone. Nothing else on this list matters if the engine is receiving degraded audio.
The usual culprits, roughly in order of how often they are the answer:
Bluetooth headsets switching to call mode. This is the most common invisible cause on both Mac and iPhone. When a Bluetooth device is used for input as well as output, it typically drops into a low-bandwidth microphone mode. Music sounds suddenly worse, which is your clue. The audio going to the transcription engine is worse too. AirPods and similar devices are excellent for listening and mediocre for dictation, for this reason and no other. We wrote separately about Mac dictation problems with AirPods.
The wrong input device is selected. Check what macOS thinks it is listening to. A monitor with a built-in mic, a webcam, or a headset you forgot was connected can be the active input while you speak into your laptop.
Distance and direction. A built-in laptop microphone is fine at normal desk distance and poor across a room. Doubling the distance costs you a lot of signal relative to room noise.
Steady background noise. A fan, an air conditioner, a nearby conversation, or a session playing through speakers all raise the noise floor. If you have music or a call playing through monitors while dictating, the mic hears that too. See fixing dictation in background noise.
The fix: use a wired input or the built-in microphone rather than a Bluetooth headset, sit at a normal distance, and remove steady noise sources where you can. If you dictate for hours a day, a modest dedicated microphone is a real upgrade, and our microphone guide for dictation covers what actually matters.
Cause 5: How you speak when you are composing
Read the test paragraph aloud and you sound confident and even. Compose an email out loud and you sound like a different person. Trailing off, restarting mid-sentence, mumbling the last three words, dropping in volume as you reach the end of a thought.
This is why some people get an excellent test result and poor real-world results. The engine is fine. The delivery changed.
Three things help:
- Think first, then speak. One clear sentence beats one sentence assembled live out of four restarts. This is a habit, and it takes about a week.
- Keep your volume up to the end. Most trailing-off happens in the final few words, which is exactly where the engine has the least following context to recover from. Ends of sentences are the weak point.
- Do not fight the pauses. If you need to think, stop dictating, think, then start a new dictation. Filling the silence with "um, so, basically" gives the engine noise to interpret.
Shorter dictations also outperform very long ones for a simple reason: a two-minute continuous take has more places to trail off, and if something does go wrong you re-record two minutes instead of two sentences.
Cause 6: Language and accent mismatch
If your output is in the wrong language, or arriving with spellings from a different variant of your language, the setting is wrong rather than the audio. This is a quick fix and easy to overlook because it can change without you touching it, particularly after a system update. Our walkthrough for the Mac side is here, and there is an iPhone version as well.
Accents are a different matter, and the honest position is this: modern transcription handles a wide range of accents far better than the systems people remember from a decade ago, but not equally. If you speak English with an accent that is less represented in general usage, you will see a somewhat higher error rate, particularly on proper nouns and technical vocabulary.
The practical response is the same personal dictionary strategy from cause 2, applied more aggressively, because the words that fail for accent reasons overlap heavily with the words that fail for vocabulary reasons. Our post on improving accuracy with an accent goes into the specifics.
Cause 7: The app is changing your text after it arrives
This one catches people out because the transcription is correct and the text on screen is not.
Between the engine producing text and you seeing it, several things can rewrite it. macOS Text Replacement entries substitute strings you set up and forgot about years ago. Autocorrect in messaging and note apps rewrites words after insertion. Autocomplete popups in editors can commit a suggestion when a newline arrives, silently swapping the word you dictated for one you did not. Grammar and writing assistants running as extensions edit text in the background.
How to confirm it: dictate the same sentence into TextEdit and into the app where you saw the problem. If TextEdit is clean and the app is not, the app is doing it. That single comparison saves an enormous amount of wasted troubleshooting.
The fix depends on the app, but the usual candidates are System Settings for Text Replacement entries, the app's own autocorrect or autocomplete preferences, and any writing-assistant extension you have installed. If the app is a third-party one behaving oddly with inserted text generally, our guide to dictation in third-party apps covers the broader pattern.
What to proofread, in priority order
Because dictation errors are real words in grammatical positions, generic proofreading is inefficient. Check these first:
- Negations. Can and cannot, will and will not, is and is not. Rare, but the most damaging single error possible, because the sentence reads perfectly and means the opposite.
- Numbers and dates. A wrong digit has no redundancy and looks exactly as correct as the right one. This is why we recommend typing numbers rather than speaking them in anything consequential, as covered in dictating numbers and dates.
- Names. Both people and products. A misspelled name is the error readers judge hardest.
- Homophone pairs. Their, there, its, it's, to, too. A skim is enough once you know to look.
- The last few words of long sentences. Where trailing off lives.
Everything else, prose will defend itself. A wrong word in ordinary prose reads oddly, and either you catch it or the reader does and understands what you meant. The dangerous errors are the ones in data, and data is the thing you should be typing.
Fixing errors without touching the keyboard
When you do find a wrong word, the repair usually costs more attention than the error did, especially on a phone where you have to place a cursor precisely between two words on glass.
Voice Keyboard Pro's iPhone keyboard includes Voice Edit for exactly this: you speak the change you want rather than performing the fine-motor task of selecting and retyping. On the Mac, the faster habit for a badly mangled sentence is usually to select the whole sentence and dictate it again rather than surgically repairing three words in the middle.
The general principle is worth internalising: dictation gets faster when you stop treating each error as a repair job and start treating a bad sentence as something to say again. Speech is cheap. That is the entire advantage.
The order to work through this
If you want the shortest path from frustrated to working:
- Run the 60-second test in a plain text field.
- If the output was mangled, fix the audio first. Nothing else matters until that is right.
- Add your twenty most-used proper nouns and specialist terms to Smart Vocabulary. This is the single biggest quality-of-life change for most people.
- Switch to speaking in complete sentences rather than fragments. Give it a week.
- Check the app you dictate into for autocorrect and text replacement interference.
- Adopt the priority proofreading list, and type numbers rather than speaking them.
Most people who conclude dictation "is not accurate enough" never did steps 2 and 3. Those two account for the great majority of real-world wrong-word errors.
Frequently asked questions
Why does dictation get common words wrong but technical words right?
Usually the opposite of what it looks like. If a common word is wrong it is generally a homophone or a boundary split, resolved by giving more context in longer sentences. Technical words come out right when they are common enough in general usage to be in the vocabulary already.
Does speaking more slowly improve accuracy?
Speaking more clearly does. Speaking more slowly, to the point of separating every word, usually makes things worse, because it strips out the context that resolves homophones and word boundaries. Aim for an unhurried conversational pace rather than a slow one.
Why is it accurate in one app and bad in another?
Because the second app is processing the text after it arrives. Autocorrect, autocomplete and text replacement are the usual causes. Dictate the same sentence into TextEdit to confirm.
Will a better microphone fix wrong words?
It fixes one of the seven causes. If your test paragraph came out mangled, yes, dramatically. If your test paragraph came out clean, a better microphone will change almost nothing, and the fix you need is a personal dictionary or a delivery change.
Can I stop it getting my colleagues' names wrong permanently?
Yes. That is precisely what Smart Vocabulary replacement rules are for. Add the name once and stop correcting it.
Is this different from the dictation built into macOS?
The seven causes are universal, since they come from how speech and language work rather than from any one product. What differs is the tools available to address them, particularly a personal dictionary with replacement rules and the ability to fix text by speaking rather than typing.
Does my audio get stored somewhere?
Not with Voice Keyboard Pro. The server stores only operational pings. No audio and no transcript content is retained. If that matters for your work, read our privacy policy rather than taking a sentence in a blog post as sufficient.
The realistic expectation
Perfect transcription does not exist, and neither does perfect typing. The comparison that matters is not dictation against a flawless ideal, it is dictation against what you actually produce with a keyboard, including the typos, the backspacing and the fact that most adults type around 40 words per minute while comfortable speech runs at 130 to 150.
Fixed audio, a personal dictionary and full sentences get most people to the point where the correction pass is short enough that the speed advantage survives comfortably. That is the goal, and it is reachable in about twenty minutes of setup rather than weeks of practice.
Voice Keyboard Pro has a free tier with daily limits, which is enough to run the test above and build a vocabulary list. Pro is $4.99 a month or $34.99 a year.