Insights & Use Cases
September 23, 2026

MacWhisper vs. Superwhisper vs. Blurt: local models, hybrid apps, and a hosted dictation API

A transcription workstation, a cross-platform dictation app, and an open-source tool built on a hosted API — three different jobs, and how to tell which one is yours.

Kelsey Foster
Growth
Reviewed by
No items found.
Table of contents

The short version: MacWhisper is the best choice for transcribing files and meetings locally, with a one-time payment and no subscription. Superwhisper is the best choice if you want one app across Mac, Windows, iOS, and Android with both local and cloud models. Blurt is the best choice if you want open source and finished text rather than a transcript, and you’re willing to send audio to the cloud to get it.

These three are compared constantly and they’re not really competing for the same job. Sorting out which job each one is for turns out to be more useful than arguing about which is better.

What each one actually is

MacWhisper is a macOS app that puts a real interface on Whisper. It transcribes audio and video files, records and transcribes meetings across the major conferencing platforms, does speaker recognition, batch processing, subtitle generation, and YouTube transcription — and it also does system-wide real-time dictation. Currently on version 11. It’s proprietary: a GUI built on open models, not an open-source app.

Superwhisper is a cross-platform dictation app for Mac, Windows, iOS, and Android. It runs local Whisper models offline on Apple Silicon and offers cloud models as well, with LLM post-processing through a range of providers and predefined Modes that reshape output for email, messages, legal writing, or casual chat. Also proprietary.

Blurt is an MIT-licensed macOS dictation app built on a hosted API. Hold the right ⌘ key, talk, finished text appears in whatever app has focus. It does one thing, and it does it in the cloud.

So: a transcription workstation, a cross-platform dictation app, and a single-purpose open-source dictation tool. Comparing them as if they’re interchangeable is where most roundups go wrong.

The real axis: where does the model run?

Everything that matters downstream follows from this one decision.

MacWhisperSuperwhisperBlurt
Local processingYesYes (Apple Silicon)No
Cloud processingYes (Pro)YesYes, only
Works offlineYesYesNo
Audio leaves your machineOnly if you choose cloudOnly if you choose cloudAlways
Open sourceNoNoYes, MIT
PlatformsmacOSMac, Windows, iOS, AndroidmacOS (Apple Silicon)
Cleanup / rewrite passVia cloud AI integrationsYes, LLM post-processingYes, in the same call
Primary jobFile & meeting transcriptionCross-platform dictationSystem-wide dictation
Pricing modelFree tier; Pro €64 one-timeFree tier; subscription + lifetimeFree app; ~$0.62/hr usage

Worth being precise about a thing that’s often stated wrongly: MacWhisper is not local-only. It runs local models by default and leans on that in its marketing, but Pro also offers cloud transcription alongside cloud AI integrations with a range of LLM providers. Superwhisper is likewise both. Blurt is the only one of the three with no local mode at all.

That makes the honest framing less tidy than “local versus cloud” and more like: two apps that let you choose per task, and one that made the choice for you in exchange for doing more with the result.

What going to the cloud actually buys

This is the part that decides whether Blurt’s tradeoff is worth it to you, so here’s the case in specifics rather than adjectives.

Accuracy on short audio. Dictation is short-form — a sentence or two at a time — and that’s the hardest case for a small local model, because there’s little context to work with. On short-form English, Universal-3.5 Pro Realtime posts a mean of 3.87% normalized word error rate across Common Voice and both LibriSpeech test sets, ranking first of the nine systems benchmarked. The full comparison and methodology are on our benchmarks hub:

SystemMean (3 sets)Common VoiceLibriSpeech cleanLibriSpeech other
AssemblyAI Universal-3.5 Pro Realtime3.87%6.44%1.88%3.28%
Speechmatics Enhanced Realtime4.69%6.78%2.31%4.99%
Smallest Pulse5.41%9.34%2.19%4.69%
Azure Realtime STT5.54%8.68%2.44%5.49%
xAI Grok Streaming5.81%11.04%2.04%4.36%
Speechmatics Standard Realtime6.41%9.43%3.18%6.61%
Deepgram Flux EN6.53%7.68%3.56%8.35%
Mistral Voxtral Mini Realtime6.61%12.23%2.10%5.49%
Deepgram Nova-37.46%12.38%3.28%6.72%

A cleanup pass in the same request. This is the larger practical difference. Blurt sends audio to the Dictation API, which transcribes and rewrites in a single call — so what arrives isn’t a transcript, it’s the message. Said: “um so can we uh move the the meeting to thursday i think friday works better actually.” Returned: “Can we move the meeting to Friday? That works better.” It resolved Thursday to Friday, which a transcript won’t do for you.

MacWhisper and Superwhisper can both reach similar output, but the path is different — a second step through a cloud LLM you’ve configured, which means more setup and, in MacWhisper’s case, your own provider keys.

Mid-sentence code-switching across 19 languages. If you switch languages inside a sentence, this is the difference between usable and not.

Speed. Around 134ms p50 on transcription, cleanup typically inside a second.

What staying local actually buys

Equally specific, because this side has the stronger argument for a lot of people.

Your audio never leaves the machine. Not “is encrypted in transit,” not “is deleted after processing” — never leaves. For clinical notes, legal work, privileged conversations, or anything under an agreement that forbids third-party processing, this isn’t a preference and no accuracy number changes it.

It works on a plane. And in a basement, and on hotel wifi that’s technically connected.

No per-hour cost and no API key. MacWhisper Pro is €64 once, with lifetime updates included. If you transcribe a lot of long audio, a one-time payment against per-hour billing isn’t close.

No dependency. The app keeps working if a vendor changes pricing, deprecates a model, or goes away. That’s worth something for a tool you use every day.

You can process files in bulk. MacWhisper’s batch transcription on local models has no metered cost, which is the whole reason journalists and researchers use it.

Compare The Numbers Yourself

Word error rate, entity error rate and latency across models and providers, with the methodology published alongside.

View benchmarks

Which is better, MacWhisper or Superwhisper?

They’re built for different jobs, so the answer depends on what you’re doing.

Choose MacWhisper if your work is files and meetings — transcribing interviews, recording calls, generating subtitles, batch-processing an archive. Its meeting integrations, speaker recognition, subtitle generation, and workflow automations have no equivalent in Superwhisper. The one-time €64 Pro license with lifetime updates is also the better economics for heavy use.

Choose Superwhisper if your work is dictation and you’re not only on a Mac. It’s the only one of the three with Windows, iOS, and Android builds, and its Modes system — switching output style between email, message, legal, and casual — is genuinely useful if you dictate into very different contexts all day.

If you do both jobs, people often run MacWhisper for files and something lighter for system-wide dictation. That’s a reasonable setup, not a failure to decide.

Where Blurt fits, and where it doesn’t

Blurt makes sense if you want open source with the cleanup pass intact, you’re on Apple Silicon, and cloud processing is acceptable for what you dictate. It’s the only MIT-licensed option of the three, which means the privacy claims are checkable rather than promised — API key in the macOS Keychain, no audio stored, no transcripts stored, no telemetry, and it explicitly refuses to read password fields or transmit your app name, window title, or selection.

Blurt doesn’t make sense if you need offline, you’re not on a Mac, you’re not on Apple Silicon, or you need file and meeting transcription. It does system-wide dictation and nothing else. Against MacWhisper’s feature list that’s a narrow tool, and deliberately so — it exists partly as a reference implementation of what the Dictation API can do.

Which leaves a gap worth naming, because it’s the one thing a three-way app comparison can’t answer for you. If your job is files rather than dictation, MacWhisper is the app and nothing here changes that — you get a window, a queue, and a transcript without writing a line of code. But if you’re processing an archive programmatically rather than clicking through it one recording at a time, that’s a different tool entirely. AssemblyAI’s pre-recorded Speech-to-Text API transcribes recorded audio at $0.21 per hour, billed per second with no minimums, with speaker diarization as a $0.02/hr add-on. It has no interface at all, which is rather the point — it’s for the case where the interface is your own script. Pick by whether you want to click or to code, not by which one is better.

There’s also a cost question worth doing honestly. At roughly $0.62 per hour of audio, a moderate dictation habit runs a few dollars a month, which beats a subscription. Heavy all-day dictation could exceed MacWhisper’s one-time €64 within a year. Work out your own hours rather than trusting either framing.

A note on published pricing

MacWhisper’s Pro license is €64 as a one-time payment on the official site — third-party roundups frequently quote €59, which is out of date. There’s a free tier, and a 25% discount for journalists, students, and non-profits by request.

Superwhisper’s pricing we could not verify reliably. Repeated checks of the official page returned inconsistent figures because the pricing widget renders dynamically, and the direct pricing URL 404s. The structure is Free / Pro (monthly or annual) / Lifetime / Enterprise, with a student discount and a refund window — but check the site yourself rather than trusting a number from any comparison post, this one included.

That’s a small thing, but it’s a useful illustration of why these roundups are unreliable in general. Most of them copy each other.

Build On The Model Underneath

Blurt is MIT-licensed and built on the Dictation API — finished text and the verbatim transcript in one call, 19 languages, $0.62/hr with everything included.

Sign up free

Pick by constraint, not by review

Your situationThe answer
Audio cannot leave your machineMacWhisper or Superwhisper on a local model
Transcribing files, interviews, or meetingsMacWhisper
You need Windows, iOS, or AndroidSuperwhisper
You want open source you can read and forkBlurt
You want finished text, not a transcriptBlurt, or Superwhisper with post-processing
You dictate across several languages mid-sentenceBlurt
You transcribe many hours and hate subscriptionsMacWhisper Pro, one-time
You’re adding dictation to your own productThe Dictation API

What the comparison is really about

The argument between local and hosted models tends to get framed as privacy versus accuracy, and that framing is about to stop being useful.

Local models are improving fast, and the gap on clean audio is already small — look at the LibriSpeech clean column above and note that several systems are inside two percentage points of each other. The remaining advantage of a hosted model shows up on the hard stuff: heavy accents, background noise, unfamiliar proper nouns, long alphanumeric strings, and code-switching. Those are exactly the cases where a bigger model has more room to reason.

But here’s the thing that’s changed in the last year, and it’s not about transcription accuracy at all. The interesting work has moved past getting the words right and into getting the text right — resolving self-corrections, stripping filler while keeping tone, applying house style, shaping output for where it’s going. That’s a language-model job, not a speech-model job, and it’s the part that’s genuinely hard to run on a laptop that’s also running everything else you’re doing.

So the durable version of this question probably isn’t “local or cloud.” It’s how much of the distance between what you said and what you meant to send you’re willing to cover yourself. MacWhisper and Superwhisper hand you excellent raw material and some tools to refine it. Blurt hands you the finished thing and asks for your audio in return.

Neither is the right answer. But knowing which question you’re answering makes the choice take about thirty seconds.

Try The Hosted Side Of The Tradeoff

Finished text and the verbatim transcript from a single request, 19 languages, $0.62/hr with every feature included. Free credits on new accounts.

Sign up free

Frequently asked questions

Which is better, MacWhisper or Superwhisper?

MacWhisper is better for transcribing files, recordings, and meetings, with speaker recognition, batch processing, subtitle generation, and a one-time €64 Pro license. Superwhisper is better for everyday dictation and is the only one of the two that runs on Windows, iOS, and Android as well as Mac. They overlap on system-wide dictation but are built around different primary jobs.

Is MacWhisper fully local?

No, though it runs local models by default and emphasizes that in its marketing. MacWhisper Pro also offers cloud transcription and cloud AI integrations with a range of LLM providers. You can use it entirely locally if you choose local models and skip the cloud integrations.

Is MacWhisper open source?

No. MacWhisper is a proprietary macOS app built on open speech models such as Whisper and Parakeet, but the application itself is closed source. Among Mac dictation tools, Blurt is MIT-licensed and VoiceInk is GPLv3 if an open-source license is a requirement.

What’s the difference between a local dictation app and a hosted dictation API?

A local app runs a speech model on your own hardware, so audio never leaves the machine and there’s no per-use cost, but model size is limited by what your Mac can run. A hosted API sends audio to a server running a larger model, which improves accuracy on accents, proper nouns, and alphanumeric strings, and can add an LLM cleanup pass that turns speech into finished text. The tradeoff is that the audio leaves your machine and you pay per hour of use.

Is there a MacWhisper alternative for transcribing files at scale?

If you want an app, MacWhisper’s own batch processing on local models is the answer and it has no metered cost. If you’re transcribing an archive programmatically, AssemblyAI’s pre-recorded Speech-to-Text API handles recorded audio at $0.21 per hour, billed per second with no minimums, with speaker diarization available as a $0.02/hr add-on. The choice is really whether you want a graphical interface or an API you call from your own code.

Does Blurt work offline?

No. Blurt sends audio over HTTPS to the AssemblyAI Dictation API for transcription and cleanup, so it requires a connection. If offline operation is a requirement, MacWhisper and Superwhisper both run local models, and Handy is a free, MIT-licensed option that works entirely offline.

How much does each one cost?

MacWhisper has a free tier and a Pro license at €64 as a one-time payment with lifetime updates included. Blurt is free and open source, billed through your own AssemblyAI API key at roughly $0.62 per hour of audio, with free credits for new accounts. Superwhisper has a free tier plus paid plans, but its published pricing was inconsistent across checks — confirm current figures on their site directly.

Title goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Button Text
Dictation