← Back to Blog

Short answer: TurboScribe transcribes audio and video files you upload. Voice Keyboard Pro is not that. It is live dictation that types at your cursor as you speak, on Mac and iPhone. If your recordings exist only because typing is slow, live dictation removes the recording step entirely.

People search for a TurboScribe alternative for two very different reasons, and the honest answer depends entirely on which one you are.

The first reason is that you have files. Interviews, lecture recordings, podcast episodes, voice memos, video footage. They already exist, they need to become text, and you want a service that does that well at a price you like. If that is you, stay in that category. Compare upload based transcription services against each other on the things that actually matter for files, and pick the winner.

The second reason is quieter, and it is the one worth thinking about. You started recording things because writing them out was slow. The voice memo you sent yourself walking to the car, the ramble you recorded because you could not face typing the email, the meeting you recorded because you knew you could not take notes fast enough. Those recordings are not the goal. They are a workaround for a bottleneck, and the transcription service is a workaround for the workaround.

This article is mostly for the second group. But let us be extremely clear about what we are and are not first, because the worst outcome is you installing something that does not do the job you came for.

What Voice Keyboard Pro does not do

Front loaded, so nobody wastes an install:

If any single item on that list is the reason you came, you can stop reading. A file transcription service is the right tool and we are not a substitute for one. For the specific case of interview recordings, we wrote up how we would actually handle it in transcribing interview recordings on Mac.

The two jobs hiding under the word "transcription"

The word covers two things that share almost nothing operationally.

Job one is archival. Audio exists. It was produced by something you do not control, like a two hour interview, a court hearing, a lecture, a podcast, a video you are editing. You need a faithful text record of it, often with timestamps and speaker attribution, sometimes as a legal or editorial artifact. The value is fidelity to a fixed source.

Job two is composition. Nothing exists yet. There is a thought in your head and it needs to become text in a document, an email, a message, a ticket, a form field. The audio is not an artifact you care about. Nobody will ever listen to it. The value is speed from thought to text.

Upload based services do job one. That is what they are architected for: ingest a file, process it, hand back a document. Live dictation does job two: you press a key, talk, and words appear where your cursor already was.

Confusing the two is easy because both convert speech to text. But the workflows are not remotely alike, and using an archival tool for a composition job is where all the friction comes from.

Count the steps

Here is a composition task run through an upload service. You want to write a paragraph of an email.

  1. Open a recording app.
  2. Record yourself saying it.
  3. Stop, save, name the file.
  4. Open the transcription service.
  5. Upload the file.
  6. Wait.
  7. Open the result.
  8. Select the text.
  9. Copy it.
  10. Switch to the email.
  11. Paste.
  12. Clean up the formatting.

Twelve steps. Now the same task with live dictation: hold a key, say the paragraph, release. The text is in the email. One step, and you never left the window.

That difference is why people who try to use a file service for daily writing quietly stop after a week. It is not that any single step is hard. It is that twelve steps of overhead for one paragraph is worse than just typing it, so you type it, and the tool goes unused.

The math that makes this worth caring about

The average adult types around 40 words per minute. Professional typists land somewhere around 80 to 100. Comfortable speech is 130 to 150 words per minute, and you already have that speed. You did not train for it.

So for anything that is fundamentally you putting your own thoughts into text, speaking is roughly three times faster than typing for most people. That gap is the entire reason live dictation exists as a category.

But the gap only pays off if the overhead around it is close to zero. Three times faster at producing words, minus twelve steps of file shuffling per paragraph, is slower than typing. Three times faster with a single keypress is just three times faster.

The recording is not the product. The text is. Any step that exists only to move audio from one place to another is pure overhead.

What replaces what

Not everything an upload service does has a live equivalent. Here is the honest mapping.

Voice memos you record to remember something

Fully replaceable. This is the clearest win. If you are recording yourself so you can transcribe it later and turn it into a note, a task, or a message, you can skip both steps and dictate directly into the destination. The memo was never the point.

Long form drafting

Fully replaceable, and better. Dictating a draft straight into your writing app means you see the text as it forms and can course correct in real time. With a recording you have to hold the whole thing in your head, then reconcile a wall of transcript against what you meant. The feedback loop is the advantage, not just the speed.

Meetings you attend

Partly replaceable. Meeting Mode on Mac captures the meeting live with speaker detection and produces AI notes, and calendar meeting detection means it can recognize when a meeting is starting. That covers most of why people record meetings. It does not cover uploading a recording that someone else made and sent you. More on how we approach it in meeting transcription on Mac.

Interviews you conduct for publication

Not replaceable. If you need a verbatim record of another person speaking, with timestamps you can cite, keep using a file based service. That is a genuine job we do not do.

Video captions and subtitles

Not replaceable. No SRT, no VTT, no time coded output of any kind. Different tool.

Lectures and archival audio

Not replaceable. Someone else's audio, already recorded, needs an upload workflow.

Roughly speaking, if the audio was produced by you for the sole purpose of becoming text, live dictation replaces the whole pipeline. If the audio is a record of something that happened, keep the file service.

How Voice Keyboard Pro actually works

On Mac

It lives in the menu bar. You hold a hotkey, speak, release, and the text appears at your cursor in whatever app is frontmost. That is system wide, so it works in your email client, your editor, a browser text field, a terminal, a form, a chat window. There is no app to switch to, which is the entire design point.

Beyond basic dictation there are three things worth knowing about:

On iPhone

It is a custom keyboard with a mic button built into it, so dictation works in any iOS app that accepts a keyboard. No sharing to a separate app, no round trip. It also does:

Privacy, since audio is involved

Any tool that handles your voice deserves a direct answer here rather than a paragraph of reassurance.

Our server stores only operational pings. No audio and no transcript content. What you dictate is not accumulating in an account somewhere waiting to be searched, exported, or breached, because there is no transcript library on our side to accumulate into.

This is a structural difference from the upload model rather than a policy difference. An upload service has to store your files and transcripts, because the stored transcript is the product you are paying for. That is not a criticism of how those services operate, it is just what the architecture requires. If you have handled anything confidential, read whichever policy applies to the tool you choose rather than taking any vendor's word for it, including ours.

Pricing

There is a free tier with daily limits, which is enough to answer the only question that matters: does dictating into your actual work feel faster than typing it. Pro is $4.99 a month or $34.99 a year.

We are deliberately not quoting TurboScribe's pricing or limits here, because those change and we would rather you check their current site than trust a number in a blog post that may be months stale. Compare the live pages.

A one week test

If you are genuinely unsure which category you need, this settles it faster than any comparison table.

For one week, every time you are about to record a voice memo, stop and ask one question: is this recording going to be listened to by anyone, ever, or is it just a slow route to text?

Keep two counts. Recordings that are artifacts, meaning someone might play them back. Recordings that are transport, meaning they exist only to be converted and discarded.

At the end of the week look at the ratio. If most of your audio is artifacts, you need a file transcription service and you should go pick a good one. If most of it is transport, you have been paying a twelve step tax on something that should take one step, and live dictation is the fix.

Plenty of people land in the middle and end up using both, which is a perfectly sensible outcome. A file service for the interviews and a dictation tool for the daily writing is not a compromise, it is just two tools doing two jobs.

Frequently asked questions

Can I upload an audio file to Voice Keyboard Pro?

No. There is no file upload, no batch queue, and no folder import. It transcribes what you speak live, at your cursor.

Does it export SRT or VTT subtitles?

No. There is no timestamped or time coded output. For captions you need a file based tool.

Does it identify different speakers?

Meeting Mode on Mac does speaker detection during a live meeting. There is no speaker separation for recordings, because there is no way to give it a recording.

Is it available on Windows or Android?

No. Mac and iPhone only. That is a real limitation, not a coming soon.

Does it work without an internet connection?

No. Transcription requires a connection.

How is this different from the meeting note tools?

Meeting assistants capture calls and summarize them. That overlaps with Meeting Mode but not with dictation, which is about writing everything that is not a meeting. We covered that distinction in our Otter alternative comparison.

What if I want both?

Use both. They cost different money for different jobs and they do not conflict.

The bottom line

If you came here with a folder of recordings that need to become documents, TurboScribe and the services like it are built for exactly that, and nothing in this article should talk you out of one. We do not upload files and we are not pretending otherwise.

But if you noticed that you are recording yourself more than you used to, and that most of those recordings get transcribed once and never opened again, the recording step is the thing to question. It exists because typing at 40 words per minute could not keep up with thinking. Talking at 130 can, and it can go straight into the document without the round trip.

Voice Keyboard Pro is free to try with daily limits. Dictate your next three emails directly into your mail app and compare it against the record, upload, wait, copy, paste loop you have been running.