← Back to Blog

Short answer: Click into Grok's prompt box, hold your dictation hotkey, speak the prompt, and release. The text lands in the box as editable text you can revise before sending. Speak the context and constraints; type the code, file paths and exact numbers.

Here is a pattern worth noticing in your own use. You open Grok with a real question, you start typing, and somewhere around the second line you shorten it. The background you were going to give, the constraint about who the output is for, the thing you already tried that did not work, all of it gets trimmed. You send eleven words instead of ninety, and then spend three follow-up messages adding back the context you removed.

That is not a prompting skill problem. It is an input speed problem. Typing runs around 40 words per minute for most adults, and 80 to 100 for genuinely fast typists. Speaking runs 130 to 150. The prompt you would have written if writing were free is a different prompt from the one you actually sent, and the answer you got reflects the one you sent.

This guide covers how to dictate prompts into Grok on a Mac and an iPhone, what to speak and what to keep on the keyboard, and the specific failure modes that make people try dictation once and give up.

The setup, in one paragraph

Dictation that works at the cursor does not need to know anything about Grok. On a Mac with Voice Keyboard Pro installed, you click into the prompt box, hold your hotkey, speak, and release. The text appears where the cursor is. Nothing gets installed into the browser, no extension asks for permission to read the page, and the behavior is identical whether you are using Grok in a browser tab, in its own app, or inside X. On iPhone, the equivalent is a custom keyboard with a microphone button, which works in any app that shows a keyboard, including the ones that make you swipe past three panels to find a text field.

If you use more than one assistant, the same key works across all of them. We have written the same workflow for dictating Claude prompts, dictating Gemini prompts and dictating in Perplexity, and the habits transfer directly.

Dictating a prompt is not the same as talking to a voice mode

Worth separating these before anything else, because they solve different problems and people conflate them constantly.

If the app you are using offers a spoken conversation mode, that is a real-time exchange. You talk, it talks back, and the experience is closer to a phone call than to writing. It is excellent when your hands are busy and you want something explained.

Dictating a prompt is the opposite shape. You are producing text that you can see, edit, reorder, shorten and reuse before anyone acts on it. The output stays on screen. You can add a line, delete a sentence that came out badly, paste in a snippet, and send when you are ready. You can copy a prompt that worked and keep it.

The practical rule: use a conversation mode when you want an answer you will listen to and forget, and dictate when you want a prompt you will edit, reuse, or hold to a standard. Most serious work is the second kind.

The shape of a spoken prompt

A prompt spoken well has the same anatomy as a prompt typed well. The difference is that speaking makes the expensive parts cheap, so you stop skipping them. Say it in beats, pausing between them:

  1. Who you want it to be. "You are reviewing this as a hiring manager who has read four hundred of these."
  2. The situation. Two or three sentences of real context. This is the part typing kills.
  3. The actual task. One sentence, stated plainly.
  4. Constraints. Length, tone, audience, what to avoid, what you have already tried.
  5. The output format you want. A table, five bullets, a draft email, a numbered plan.

Read that list back and notice which parts you routinely omit when typing. For most people it is two and four, which are also the two that most determine whether the answer is useful. A generic answer is very often the correct response to a prompt with no context and no constraints.

The context dump you would never type

This is the single highest-value thing dictation unlocks. Before asking the question, spend forty seconds describing the situation as if to a colleague who just walked in. What the project is. Who it is for. What the deadline pressure is. What you already ruled out and why. What the last attempt produced and what was wrong with it.

Forty seconds of speech is roughly a hundred words. Typing a hundred words of context costs you two and a half minutes, which is exactly why nobody does it. Speaking it costs almost nothing, and it changes the answer more than any clever prompt phrasing.

Think out loud, then cut

Do not try to speak a polished prompt on the first pass. Dictate the messy version, including the false starts, then read what is on screen and delete what does not earn its place. Editing text you can see is much easier than composing perfect sentences in your head while the microphone is open. Our post on why voice typing produces better first drafts covers the underlying reason: separating generation from evaluation makes both faster.

Dictate the follow-up instead of starting over

The common failure after a mediocre answer is to rewrite the whole prompt from scratch. Instead, speak a follow-up that says what specifically was wrong. "That is too formal and it assumes the reader knows the background. Rewrite it for someone seeing this for the first time, half the length, and drop the last paragraph entirely." Thirty words spoken, and the correction is more precise than anything you would have bothered to type.

Speak the prose, type the syntax

This is the boundary that separates people who dictate happily from people who quit in week one. Some content belongs on the keyboard, and the reason is not accuracy in general. It is that certain errors are silent.

A sentence that comes out slightly wrong reads as wrong. You see it and fix it. But a wrong digit in a figure, a mangled file path, or a slightly-off variable name looks completely correct on screen, gets sent, and produces an answer built on a false premise. You then debug the answer instead of the input.

Keep these on the keyboard:

The natural rhythm becomes: speak the framing and the question, stop, paste or type the exact material, then speak the constraints. You are not choosing between voice and keyboard. You are using each for what it is good at.

Getting the words you actually use

The complaint that ends most dictation experiments is that the tool does not know your vocabulary. Product names, internal project names, client names, technical jargon and industry acronyms come out as the nearest common English word, and after correcting the same term nine times you conclude the whole thing is unreliable.

Smart Vocabulary is the fix. You add the terms you use along with replacement rules, so the words that keep coming out wrong stop coming out wrong. For prompting specifically, the list worth building is short and very high-leverage: the names of the products and codebases you ask about, the acronyms of your field, and the names of the people and companies you keep referencing. Twenty entries covers a surprising share of the problem.

For the ones that slip through, Voice Edit lets you speak a change rather than reaching for the mouse. Highlight the sentence, say what is wrong with it, and it gets fixed. That matters more in a prompt box than in a document, because a prompt box is a bad text editor with no formatting controls and a send key waiting to be hit by accident.

Prompting in another language

If you think in one language and prompt in another, two-way translation while dictating removes the intermediate step entirely. Speak in your first language, get the text in the target language, send it. Twenty-four languages are supported.

The honest caveat: machine translation is a first draft. For a prompt, that is usually fine, because a prompt is instructions to a machine rather than a message to a person, and small awkwardness costs nothing. For anything the output of which goes to a human, read it before you send it.

The traps, in the order people hit them

The send key

The prompt box in most assistants sends on Enter. If your dictation adds a line break where you wanted a paragraph, you have just sent half a prompt. Two defenses: build the habit of composing longer prompts in a scratch space and pasting them in, and check what your dictation does with a spoken "new paragraph" before you rely on it. Our post on dictating line breaks and paragraphs goes through it.

Losing focus in the text box

Dictation types at the cursor. If the cursor is not in the prompt box, the text goes somewhere else or nowhere at all. Click into the field first, every time, and glance at it before you start speaking. This is the number one cause of the "it did not work" report, and it is a habit rather than a setting.

Confusing the assistant box with the post box

If you use Grok inside X, there are two text fields on screen with similar visual weight, and one of them publishes to the world. Look before you speak, and treat the post composer with the caution it deserves. We wrote separately about dictating posts on X if that is the workflow you want.

Speaking a wall of text with no structure

Because speech is cheap, it is easy to produce four hundred words of unpunctuated stream. That is a worse prompt than a tight ninety-word one. Pause between the beats listed earlier, say "new paragraph" between sections, and skim what is on screen before sending. The point of dictation is to lower the cost of including things that matter, not to raise the word count.

Background noise and the microphone you are using

Accuracy tracks input quality more than anything else. A laptop microphone across a desk in a noisy room will underperform a headset every time, and a Bluetooth headset that has switched to its call profile can sound noticeably worse than the same headset on a wired connection. If your results are inconsistent rather than uniformly bad, that is usually the reason.

Where this actually pays off

Four situations where dictating prompts changes the outcome rather than just the typing time:

Debugging something you do not understand yet. Describe what you did, what you expected, what happened instead, what you have already ruled out. That description is long, tedious to type, and precisely what determines whether the answer is useful. Speaking it takes under a minute.

Asking for a draft in your own voice. Instead of describing the tone you want in the abstract, speak a paragraph the way you would actually say it and then ask for the rest to match. You have supplied a sample rather than an adjective, which works considerably better than "make it friendly but professional."

Research questions with real constraints. Budget, timeline, location, what you have already tried, what is non-negotiable. Six constraints spoken in twenty seconds narrows an answer far more than a well-phrased single sentence.

Ideas that arrive away from the desk. On the phone, the keyboard's microphone button means a prompt you thought of on a walk gets captured as a full paragraph rather than a three-word note you cannot decode later.

The prompt you would write if writing were free is a different prompt from the one you actually send. Dictation makes them the same prompt.

Honest limits

Voice Keyboard Pro runs on macOS and iPhone, not Windows or Android. It needs a connection to transcribe. It does not read the page you are on, does not integrate with Grok or any other assistant, and does not know what is in your prompt box; it puts text where your cursor is, which is exactly why it works the same everywhere. On privacy, the specific claim is narrow and checkable: our server stores only operational pings, no audio and no transcript content. What the assistant you are prompting does with the text you send it is governed by its own policy, which is worth reading if the material is sensitive.

There is a free tier with daily limits. Pro is $4.99 a month or $34.99 a year.

Frequently asked questions

Does this work with Grok in a browser, or only in an app?

Both. Dictation types at the cursor at the system level, so a browser text field and an app text field behave identically. Nothing is installed into either one.

Can I dictate code into a prompt?

You can, but do not. Paste code. Speaking punctuation-dense syntax is slow, error-prone, and the errors are the invisible kind. Speak the explanation around the code instead, which is the part that is actually hard to type and easy to omit.

Will it capitalise and punctuate the prompt for me?

Yes, punctuation and capitalisation are handled as you speak. For a prompt, this matters less than it does for an email, but it makes the text much easier to reread before sending.

What if I use several assistants?

One hotkey covers all of them, plus your email, your notes app and your terminal. That is the argument for a system-level tool over per-app voice buttons: you learn one habit rather than six.

Does dictating make my prompts too long?

It can, if you never reread them. The discipline is to speak freely and then cut, which takes about fifteen seconds and produces a better prompt than either a rushed one-liner or an unedited monologue.

Can I dictate in a language other than English?

Yes, and you can also speak one language and have the text arrive in another using two-way translation across 24 languages. Check the result before sending anything that a person will read.

Try it on the next real question

Not a test prompt. The next genuine one you were about to type. Click into the box, hold the key, and say the whole thing out loud including the background you would normally cut. Then read it back, delete a third of it, and send.

The difference will not be in how fast you typed. It will be in the answer. Voice Keyboard Pro has a free tier, so the experiment costs a download and one honest prompt.