Short answer: Dictate SOAP notes section by section, not as one block. Paste your S/O/A/P headers first, click into each section, and speak that part alone. Type or click structured data like vitals and codes, add clinical terms to a custom dictionary, and reserve dictation for the narrative parts.
SOAP notes are the single most repeated writing task in clinical work, and for most people they are also the slowest. The visit takes fifteen minutes. The note takes eight. Multiply that across a full panel and the arithmetic gets ugly: the documentation is not a side effect of the day, it is a second shift stacked on top of it, usually finished after hours in a parking lot or on a couch.
Dictation is the obvious lever, and it is a good one. But most people try it once, speak a whole note into a box, get back a wall of run-on text with the wrong drug spelling and no section breaks, and conclude that voice does not work for clinical documentation. That conclusion is wrong, and the reason is almost always method rather than technology. SOAP notes have a fixed skeleton, and that skeleton is exactly what makes them well suited to voice if you dictate into the structure instead of trying to dictate the structure itself.
This guide covers the workflow: how to set it up, what to say in each section, which parts you should never dictate freehand, how to handle the vocabulary problem, and a realistic ramp so you are not slower in week one than you were before.
Why SOAP notes suit dictation better than most writing
Most writing tasks are slow because you do not know what you are going to say. Dictation does not help much there, because the bottleneck is thinking, not typing.
SOAP notes are the opposite. By the time you sit down to write, you already know the content completely. You just took the history. You just did the exam. You already formed the assessment in the room. The note is not a thinking task, it is a transcription task, and you are transcribing from your own short-term memory into a keyboard at roughly 40 words per minute while your mind runs at conversational speed.
That gap is the whole opportunity. Ordinary speech runs 130 to 150 words per minute. Even fast professional typists top out around 80 to 100 WPM, and clinical typing is slower than prose typing because of the constant stopping for terms, abbreviations, and formatting. When the content is already fully formed in your head, speaking it is two to three times faster than typing it, and that ratio holds up in practice as long as you do not fight the tool.
There is a second, less obvious benefit. Notes that get typed tend to get compressed, because typing is expensive and the writer unconsciously trims. Assessments become one line. Plans become three abbreviations. Dictated notes are usually more complete, not because anyone decided to write more, but because the cost of another sentence dropped to almost nothing. If your notes have been getting thinner over the years, input cost is a likelier explanation than clinical laziness.
The template-first method
The single biggest mistake is dictating the note as one continuous block and hoping the structure survives. It will not. Dictation engines produce a stream of text. They do not know that "objective" is a heading rather than a word you happened to say.
Instead, put the structure down first, then fill it with voice.
- Create or open your note template with the four headers already in place: Subjective, Objective, Assessment, Plan. Most systems have a template feature. If yours does not, keep a plain text snippet you paste in.
- Click into the Subjective field. Only that field.
- Hold your dictation hotkey and speak that section alone. Release. Text lands where the cursor is.
- Move to the next field and repeat.
This sounds trivially simple, and it is, which is why it works. Four short dictations of twenty to sixty seconds each are dramatically more accurate than one three-minute monologue, for a few reasons. Short passes give the engine tighter context, so a term is more likely to be interpreted within the right clinical frame. Errors stay local, so a garbled phrase costs you one section rather than forcing you to re-read the whole note. And you get natural checkpoints where you can glance at what landed before moving on.
The setup matters here. If your dictation tool requires you to open a separate window, speak, then copy and paste into the field, the template-first method falls apart because the friction per section is too high. You want system-wide dictation that types at the cursor in whatever field is focused, so switching sections costs one click. On a Mac, Voice Keyboard Pro works this way: hold a hotkey anywhere, speak, release, and the text appears at the cursor in whatever application is in front, including a browser-based charting system.
What to say in each section
Subjective
This is the section dictation wins by the widest margin, because it is pure narrative and it is closest to how you would describe the visit out loud to a colleague.
Speak it as reported speech, in complete sentences, in the order you took the history: chief complaint, history of present illness, then the relevant review of systems. Do not try to compress while you speak. Compression is an editing operation, and editing while dictating is the fastest way to produce a stalled, fragmentary note.
A useful habit is to speak the patient's own framing first and your clinical translation second, since that is the distinction the section exists to preserve. If the patient said the pain is like a band around the head, dictate that phrase. You lose the texture if you jump straight to the label.
Objective
This is the section where you should dictate least. Objective data is structured data: vitals, measurements, lab values, exam findings with laterality and grading. Numbers are the weakest thing to send through any speech engine, not because engines cannot hear digits, but because the failure mode is silent. A misheard word looks wrong on the page. A misheard number looks perfectly normal and reads as a real value.
The rule: pull vitals and results from the source, click or type them, and dictate only the narrative parts of the physical exam. "Lungs clear to auscultation bilaterally, no wheezes or crackles" is a great dictation. "Blood pressure one thirty two over eighty four" is a bad one, not because it will probably fail, but because when it does fail you will not notice.
If you do dictate numbers, read them back before moving on. This is the one place in the note where a proofread pass is genuinely non-optional. We have written more about this pattern in dictation tips for better accuracy, and the guidance is the same in every domain where numbers carry weight.
Assessment
The assessment is where dictated notes tend to be noticeably better than typed ones, because reasoning is expensive to type and cheap to say.
Dictate your actual clinical reasoning, not just the label. Speak the working diagnosis, then the differential you considered and why you moved away from it, then the degree of certainty. This is the part of the note that has real downstream value, for the next clinician, for continuity, and for you in six months when the patient returns and you cannot remember what you were thinking. It is also the part most often reduced to a stub because it takes too long to type.
Diagnosis codes are a separate problem. Speaking an ICD-10 code aloud is a poor use of voice. Use your system's code picker, and see medical dictation with drug names and ICD-10 codes for a fuller treatment of why coded fields and dictation should stay separate.
Plan
Plans are lists, and lists dictate well if you speak the list markers explicitly.
Say "number one" or "next line" between items rather than expecting the engine to infer structure from your pauses. It cannot. Dictate the medication decisions, the follow-up interval, the patient education you gave, and the return precautions. Return precautions in particular are frequently truncated in typed notes and rarely truncated in dictated ones, which is a small but real safety improvement.
Fixing the vocabulary problem
This is the objection that stops most clinicians, and it is a legitimate one. General-purpose transcription is trained on general-purpose language. Drug names, anatomical terms, procedure names, and the specific surnames on your panel are all outside that distribution. You will get "metformin" reliably. You will get less common agents, brand names, and hyphenated surnames wrong often enough to be annoying.
The fix is not a better microphone and it is not speaking more slowly. Those help with acoustics, and this is not an acoustic problem. The words are being heard fine and mapped to the wrong output because the correct output is rare in general text.
What actually fixes it is a custom dictionary. Voice Keyboard Pro calls this Smart Vocabulary: you add the terms you use, with replacement rules, and they get applied to your transcriptions. Build it incrementally rather than trying to load a whole formulary on day one. Every time you correct the same word twice, add it. Within two weeks your personal list covers most of what you actually say, which is a much smaller set than the full medical lexicon.
Three categories are worth seeding deliberately:
- Your prescribing set. Not every drug, just the thirty or forty you actually write. Include the spellings you prefer, generic or brand, so you are not correcting a consistent mismatch every time.
- Local proper nouns. Referral practices, the imaging center you use, colleague names, facility names. These are almost never in a general vocabulary and they recur constantly.
- Your abbreviation preferences. If you want "shortness of breath" to stay expanded, leave it. If you want it abbreviated, set the rule once instead of editing it forever.
Related reading: how custom vocabulary learns your words goes deeper into how replacement rules compound over time.
Where to dictate: at the desk or between rooms
There are two workable rhythms and they suit different practices.
In-room or immediately after, at the workstation. Best for accuracy and completeness, because recall is perfect. The constraint is the room itself. Dictating in front of a patient changes the visit, and some patients find it clarifying while others find it clinical and cold. If you do it, narrate what you are doing first.
Between rooms, on a phone. Best for throughput. You capture the note in the corridor while it is fresh, then paste and tidy at the workstation later. This is where a phone-side voice keyboard matters: because it works at the keyboard layer, you can dictate into whatever app you keep notes in rather than being locked into one vendor's capture screen. Voice Edit is useful here too, since you can speak a correction rather than pecking at a phone screen with one hand.
The second pattern only works if capture is genuinely instant. If it takes four taps to start recording, you will stop doing it by Thursday.
The privacy and compliance boundary
This section is the one to read carefully, because it is the one that carries real consequences.
Voice Keyboard Pro's servers store only operational pings. No audio and no transcript content is retained on our servers. That is a genuine architectural fact and it is a meaningful difference from tools that keep recordings for model improvement, but it is not the same thing as a compliance certification, and we will not tell you it is.
If you are documenting protected health information, the question is not what a marketing page says. It is what your organization's policy and your business associate agreements permit for any tool that processes patient data. That determination belongs to your compliance officer, not to a blog post. Ask before you dictate PHI into any tool, including this one. We have written a longer piece on HIPAA considerations for clinical note dictation on Mac that lays out the questions worth asking, and the honest summary is that the answer depends on your setting.
Two practical habits regardless of tooling: do not dictate identifiers you do not need in the note, and be aware of who can hear you. The acoustic leak in a shared workspace is a more common real-world exposure than anything happening in transit.
Common mistakes
- Dictating the headers. Say the content, not "Subjective colon". Put headers in the template.
- Editing mid-sentence. Speaking, stopping, correcting a word, and resuming produces a worse note than speaking through and cleaning up once at the end. Let the paragraph land, then fix it.
- Dictating a whole note in one pass. Covered above, but it is the most common failure and worth repeating.
- Trusting numbers. Vitals, doses, and dates need eyes on them before you sign.
- Speaking too quietly. People instinctively drop to a near-silent murmur when documenting near others. Quiet is fine, mumbled is not. Consistent articulation at low volume works; trailing off at the end of sentences does not.
- Not adding corrections to the dictionary. If you have corrected the same term five times, that is five times you chose to keep the problem.
A seven-day ramp
Do not convert your whole documentation workflow on a Monday. The first day of dictation is slower than typing, always, and if that first day is a full clinic day you will quit.
Days one and two. Dictate only the Subjective section, and only for the last three patients of the day when the pressure is off. Get comfortable with the hotkey and with speaking in complete sentences.
Days three and four. Add the Assessment section. This is the second-easiest and the highest value. Start your vocabulary list with every term you correct.
Day five. Add Plan, with explicit list markers. Keep Objective typed.
Days six and seven. Run full notes on a normal patient load. Time yourself honestly against your typed baseline. Most people are at parity by the end of the first week and meaningfully ahead by the end of the third, once the vocabulary list has absorbed their actual working language.
If you are not faster by week three, the usual cause is one of two things: you are still dictating in one block, or you are editing while you speak. Both are habits, and both are fixable.
The realistic bottom line
Dictation will not make SOAP notes enjoyable. It makes them shorter in wall-clock time, and it tends to make them better in content, because the parts that carry the most clinical value are exactly the parts that typing punishes most.
The gains are real but they are bounded. Structured data still needs to be entered as structured data. Coded fields still need pickers. Signing still needs a read. What changes is the narrative layer, which is where most of the minutes actually go, and cutting that layer by half across a full panel is the difference between finishing notes in clinic and finishing them at nine at night.
The note is not a thinking task by the time you sit down to write it. It is a transcription task, and you are transcribing at a third of the speed you can speak.
If you want to try the workflow, Voice Keyboard Pro has a free tier with daily limits, and Pro is $4.99 per month or $34.99 per year. Start with the Subjective section on three notes tomorrow, add every corrected term to your vocabulary list, and judge it at the end of week one rather than the end of day one. For other clinical contexts, see voice to text for doctors on Mac and dictation for therapists writing session notes.