Short answer: Sonix transcribes recordings you upload. Voice Keyboard Pro types your words straight into whatever app you are already in. If you need transcripts of existing audio files, keep a transcription platform. If most of your day is writing from scratch, that is a different tool.
People search for a Sonix alternative for two very different reasons, and the advice for each is almost opposite.
The first group wants the same thing for less money, or with a different language mix, or with an export their workflow needs. They have recordings, and they need those recordings turned into text. Everything below about our app will be a poor fit for them, and we will say so plainly rather than waste their afternoon.
The second group arrived at a transcription platform for a reason that turned out to be indirect. They were slow at getting words out. Someone said "just record it and get it transcribed," so they did. Then they discovered that a transcript is not a document, and that they were now doing two jobs instead of one. That group is looking for the wrong category of tool, and this post is written for them.
The honest part first: what we do not do
Putting this at the top rather than burying it in a footnote, because it decides the question for a lot of readers.
Voice Keyboard Pro does not:
- Accept audio or video file uploads. There is no place to drop an MP3 from last week's interview. We transcribe what you speak, live, into the cursor. We do not process recordings that already exist.
- Provide a multi-speaker transcript editor. No timestamped browser workspace where you scrub audio and correct a transcript line by line.
- Export subtitles or caption files. No SRT or VTT deliverables.
- Host a shared team transcript library. There is no searchable archive of everything your organization has recorded.
- Offer a batch transcription API. If you have a pipeline that sends files programmatically and pulls back structured transcripts, we are not a drop-in replacement for it.
If any of those five is the reason you use Sonix, you should keep using a transcription platform. Nothing in the rest of this post changes that, and swapping tools would cost you a capability with no compensating gain.
Two different jobs that both involve talking
The confusion is structural, not anyone's fault. Both categories convert speech to text, so they look adjacent. They are not.
The archive job: a recording exists. An interview, a lecture, a deposition, a focus group, a podcast episode. Someone already said the words, and you need a record of them. The output is a document about a past event. This is what transcription platforms are built for, and they are good at it.
The input job: nothing exists yet. There is a blank field and a blinking cursor, and the words have to come out of your head for the first time. An email, a report, a client update, a set of notes, a chapter, a message. The output is new text that did not exist in any form five minutes ago.
Most professional writing is the input job. Count your own week: the hours you spent in recorded conversations are almost certainly a small fraction of the hours you spent producing text, and the recorded hours mostly produce reference material rather than finished work.
A transcript library is a record of what was said. It is not the thing you were supposed to write.
The friction that sends people looking
Here is the workflow that eventually feels wrong, and it is the same story from researchers, marketers, and consultants alike.
You record yourself talking through a report. Fifteen minutes of good thinking. You upload it, wait for processing, open the transcript, and find a wall of text that is exactly what you said. Which means it contains your false starts, the sentence you abandoned halfway, the part where you doubled back to re-explain something, the "um," the tangent about the other project, and no paragraph breaks that correspond to your actual argument.
Now you copy that into a document and start cutting. The cutting takes longer than the talking did. And the finished report is still not written, because a spoken explanation and a written one have different shapes, and you have just spent thirty minutes converting one into the other by hand.
That round trip has real steps: record, upload, wait, open, correct, copy, paste, restructure, rewrite. Dictating at the cursor has one: speak into the field where the text belongs. The difference is not accuracy. It is that one workflow produces raw material and the other produces text in the place it was meant to go.
What dictating at the cursor actually looks like
Voice Keyboard Pro is a menu bar app on the Mac and a keyboard on the iPhone. On the Mac you hold a hotkey, speak, and release; the text appears at your cursor in whatever application is in front of you. Mail, a browser text box, a document, a CRM field, a code comment, a chat window. The app does not care which, because it types where the cursor is.
Because the text lands in the destination, you never manage a file, never wait for a job to finish, and never move text between two applications. And because it is available everywhere on the system, it covers the small writing that never justified opening a transcription tool: the four-line reply, the two-sentence status update, the comment on a document. In most jobs, the small writing is where the hours actually go.
The speed argument is simple arithmetic. An average adult types around 40 words per minute; professional typists reach 80 to 100. Everyone speaks at 130 to 150. That gap has always existed, but it only pays off if speaking puts the words where you need them without a conversion step in between.
On the iPhone the same idea takes the form of a keyboard with a mic button, which works in any app: Slack, Notes, your email client, a CRM's mobile app. Voice Edit lets you speak a correction rather than trying to place a cursor precisely on a phone screen, which is the single most annoying part of writing anything long on mobile.
Languages: two different questions
Multilingual work is a common reason to be on a transcription platform, and the two jobs split here too.
"I have a recording in another language and need it transcribed or translated." That is the archive job. Keep a transcription platform for it. We do not process files.
"I need to write to someone in another language." That is the input job, and it is what two-way translation while dictating is for. You speak in your language, and the text arrives in theirs, across 24 languages. It works in both directions during a conversation, which makes it useful for support threads and client email rather than just one-off documents.
The caveat we would give anyone: machine translation is a first draft. It is genuinely good for a status update or a scheduling message. Have a fluent human check anything contractual, anything about money, and anything where a subtle wrong shade would be expensive. Our post on voice dictation for translators covers where the line sits in practice.
The privacy difference is structural
This one is worth understanding properly rather than as a slogan.
For any platform whose product is a transcript library, storing your content is not a side effect. It is the product. The searchable archive, the shareable link, the ability to come back to a recording from March, the team workspace: all of it requires that the audio and the transcript be kept. That is a reasonable design for the job, and reputable platforms document their retention and security posture. If you are on one, read that documentation, because for regulated work it is the document that matters.
Our design does not need to keep anything, so it does not. Our servers store only operational pings, not audio and not the content of what you dictated. Text lands at your cursor and stays in your own document, on your own machine.
We would rather you verify that than take our word for it, so read our policy yourself before you use this for sensitive material, and check whatever rules govern your own work. A privacy claim you did not check is not a privacy control.
The names problem, and the fix
Every profession has words no general transcription has encountered: client names, product lines, internal abbreviations, technical vocabulary, the surname of the person you email daily. They are a small share of your words and appear in nearly every piece of writing you produce, which is why a miss on them feels so much worse than the error rate suggests.
A transcript editor solves this by making you fix it in the editor. Smart Vocabulary solves it upstream: a personal dictionary with replacement rules, so the spellings that matter to you come out right the first time and stay right. Add fifteen terms at the start, add to it when something new appears, and the corrections mostly stop.
If what you actually wanted was meeting notes
A fair number of people land on transcription platforms because of meetings, so this is worth naming directly.
Meeting Mode on the Mac detects speakers and produces AI notes, and calendar meeting detection means it can be ready when the meeting is instead of after you remember. No bot joins the call as a visible participant, which matters more than people expect in client and external meetings.
Two honest boundaries. Those notes are raw material, not a deliverable: read them and rewrite them into whatever you are actually sending. And tell people in the room. Norms and rules on recording and note-taking vary by jurisdiction and by employer, and that call is yours to make, not ours. Our guide to meeting transcription on the Mac and our Otter alternative comparison both go further into where this fits.
Five good reasons to keep Sonix
We would rather you make the right call than switch and regret it. Keep a transcription platform if:
- The recordings are the work. Interviews, oral histories, research sessions, lectures. If the source audio is the primary artifact, you need something that processes files.
- You need a searchable shared archive. Teams that go back to old recordings need the library. We do not provide one.
- You deliver subtitles or captions. Caption files are a real output format we do not produce.
- You need verbatim records for compliance or legal purposes. A live dictation tool is not a record of a proceeding.
- You have an automated pipeline. If files arrive programmatically and transcripts feed downstream systems, replacing that with a Mac keyboard tool is not a swap, it is a demolition.
Also relevant: our Mac app is macOS only, and the keyboard is iOS. If your team is on Windows or Android, a browser-based platform will cover people we cannot.
The middle path most people end up on
The realistic answer for a lot of readers is not "switch." It is "these are two tools for two jobs, and you were using one of them for both."
Keep the transcription platform for recordings. Add a dictation tool for everything you write from scratch. The transcription bill probably shrinks, because you stop recording yourself talking through documents just to get a draft, and that self-recording is often a surprising share of the minutes.
We have written the same comparison for adjacent tools if your shortlist has more than one name on it: Trint and Notta both sit in the upload-and-transcribe category and the reasoning carries across.
A test that gives you a real answer in five days
Do not decide from a feature table. Decide from your own week.
- Day one: keep a tally in a note. Every time you produce text, mark whether it came from a recording that already existed or from a blank field. Two columns.
- Days two and three: keep working normally, keep tallying.
- Day four: install a dictation tool and use it for every item in the blank-field column. Nothing else changes.
- Day five: look at the two columns and at how day four went.
If your blank-field column is short and your recordings column is long, you are in the archive business and you should keep the platform you have. If the blank-field column dwarfs the other one, which is the usual result, then the tool that was going to help you most was never the transcription platform.
Frequently asked questions
Is it cheaper than a transcription platform?
There is a free tier with daily limits, and Pro is $4.99 a month or $34.99 a year. We are not going to compare that to anyone else's pricing, because pricing changes and we would rather you check the source. The more useful point is that it is not the same product, so the comparison is not apples to apples.
Can I upload an old recording just this once?
No. There is no upload path at all. This is the clearest line between the two categories.
Does it work in Google Docs, Word, and my CRM?
Yes. It types at the cursor system-wide on the Mac, so any application with a text field works, including web apps in the browser. On iPhone, the keyboard works in any app that accepts a keyboard.
Does it work offline?
No. Transcription runs on our cloud infrastructure and needs a connection. On a plane, macOS's built-in dictation is a reasonable fallback for short notes.
What about accents?
Accented English is handled well by modern transcription generally. The errors that will actually annoy you are proper nouns and jargon, which is a different problem with a different fix: add them to Smart Vocabulary.
Can I use it on Windows?
The desktop app is macOS only. The iPhone keyboard is iOS only.
The short version
If you have recordings and need transcripts, use a transcription platform. That is what they are for, and no dictation tool replaces one.
If the reason you started recording yourself was that writing was slow, the transcript was never the fix. It added a conversion step to a problem that was really about input speed. Speaking directly into the field where the text belongs removes the conversion entirely.
Voice Keyboard Pro has a free tier, which is enough to run the five-day test above and find out which column your week actually lives in.