DanDin

How we test dictation apps

By Published Last reviewed

The short answer

We make DanDin, a dictation app, so here is how we test dictation apps, including what we have not done yet. Today every fact about another app on our site comes from that company's own pages, with the link and the date we read it. We have not run hands-on tests of other apps, and each one is marked "not tested yet". When we test, every app will get the same script: sentences in Modern Standard Arabic, Arabic with English words mixed in, English, and dialect sentences written and read by native speakers, on one machine on one day. We score word error rate: the words that were changed, dropped or added, divided by the words in the correct text. We report it twice, once after evening out spelling details like vowel marks and letter forms, and once without, so a spelling habit does not look like a hearing mistake. We never publish scores from synthetic voices. Found an error? Email support@dandin.com.

Have you tested the apps you compare?

Not yet, and no page on this site says otherwise. What we have done is read each company's own documentation, quote it, link it, and write down the day we read it. That is enough to answer questions like which languages a company lists, what an app costs and which computers it runs on, because the company is the authority on all three. It is not enough to tell you which app hears your dialect better. Until we run the script below, every app on this site carries the words "not tested yet".

What we know about each app today, and from where. Nothing here is a hands-on test result.
AppWhat we know, and from whereHands-on test by us
DanDin (ours)Its own code, and captures made on a Mac with a synthetic voiceExamples only, no accuracy results
Wispr FlowIts help centre and pricing page, read 2026-10-08Not tested yet
SuperwhisperIts home page, Windows page and download page, read 2026-10-08Not tested yet
FreeFlowIts GitHub repository and readme, read 2026-10-08Not tested yet
Dragon Professional and Dragon LegalNuance product pages, read 2026-10-08Not tested yet
Windows voice typing and Voice accessMicrosoft help pages, read 2026-10-08Not tested yet
Dictate in Microsoft WordMicrosoft help page, read 2026-10-08Not tested yet
Apple DictationApple help and feature availability pages, read 2026-10-08Not tested yet
Google Docs voice typing and GboardGoogle help pages, read 2026-10-08Not tested yet

Swipe the table to see every column

What do we measure?

Two numbers, and one habit.

The first number is word error rate. It counts whole words: a word the app changed, a word it dropped and a word it added each count as one error, and the total is divided by the number of words in the correct text. The second is character error rate, the same idea counted letter by letter, which is useful in Arabic because a single missing letter at the end of a word can change the meaning while leaving the word looking almost right.

We also split the errors into the three kinds, because they mean different things. Changed words mean the app misheard you. Dropped words often mean it swallowed the end of a word, which is what usually goes wrong with spoken dialect. Added words usually mean it invented something out of silence or noise.

The habit we check by reading, not by counting: in a sentence that mixes Arabic and English, did the English words stay in Latin letters, the way you said them?

What we do not measure yet: speed, battery use, and anything about privacy. We have no honest way to measure those on someone else's machine, so we will not score them.

How is word error rate calculated?

Take a sentence of ten words. The app gets one word wrong and drops one word. That is two errors out of ten, which is 20 percent. Lower is better, and 0 percent means every word matched.

Here is the sentence we used for the number in the table, in English: "Please send the monthly report before Thursday with a summary." An app that typed "Please send the monthly report before Tuesday with summary" changed one word, Thursday, and dropped one word, a. Ten words in the correct text, two errors, 20 percent.

Before we compare two texts we tidy both the same way, and we do it twice because the two ways answer different questions.

  • Evened out, which our scorer calls STANDARD: remove vowel marks and the stretching character, remove punctuation, treat the alef forms and the yaa forms and the taa marbuta as the same letter, turn Arabic numerals into 0 to 9, and lowercase Latin letters. This is the headline number, because a writing habit is not a hearing mistake.
  • As written, which our scorer calls STRICT: remove only extra spaces and punctuation, and keep the spelling exactly as it came out. This is where you can see the small errors the first number hides.

An example of the difference, in Arabic. The correct text ends in صباحًا, with the vowel mark. An app that typed صباحا, without it, scores 0 percent evened out and 10 percent as written. Nothing was misheard, so the first number is the fair one, and the second number tells you the app writes vowel marks differently from the person who wrote the script.

Worked examples, computed with our own scorer on 2026-09-17. Illustrations of the arithmetic, not results for any app.
ExampleWords in the correct textChangedDroppedAddedError rate
One wrong word and one dropped word1011020%
Only a vowel mark differs, evened out100000%
Only a vowel mark differs, as written1010010%

Swipe the table to see every column

One more rule: we add up the errors and the words across all the sentences and divide once, instead of averaging the sentences. Otherwise one short sentence that went badly would count as much as a long one that went well.

The part of our tools that compares two texts never calls a model of any kind. It reads the correct text, reads what the app typed, and counts. You can check the arithmetic on this page with a pen.

What script do testers read?

One fixed script, the same for every app. It has four kinds of sentence: Modern Standard Arabic, Arabic with English words in the middle of it, English, and dialect. Here are the sentences we have written so far. The dialect lines are in preparation, and they will be written and read by people who speak those dialects.

Modern Standard Arabic:

  • أرجو إرسال التقرير الشهري قبل يوم الخميس، مع ملخص قصير لأهم النتائج.
  • ذكرني أن أشتري هدية لأمي قبل نهاية الأسبوع.
  • الاجتماع القادم سيكون في مقر الشركة الساعة العاشرة صباحًا.

Arabic with English words in it:

  • عندي meeting على Zoom الساعة ثلاثة، ذكرني أرسل الرابط في Slack.
  • خلصت الـ presentation، بأرسلها لك بعد الـ meeting.

English:

  • Please send the monthly report before Thursday with a short summary.
  • Remind me to call the bank tomorrow morning.

Who reads the script, and on what microphone?

Today, nobody. The only voice examples on this site are made by a computer voice on one machine: an iMac with an Apple M1 chip running macOS 15.6.1, using the microphone built into it. That machine has exactly one Arabic computer voice, called Majed, and it speaks general Arabic with no dialect.

Recordings by people are planned, with written permission from every speaker, and every published recording will say which machine, which microphone and which day. We will not publish a score from a recording until the person who made it has agreed in writing.

Why no scores from synthetic voices or restricted data?

A computer voice does not hesitate, does not run words together, has no dialect and sits in a silent room. An app can do well on it and still fail on you. So we use the computer voice for one thing only, to show you what DanDin types, and we write "synthetic voice" under every such example.

The same rule covers data we are not allowed to publish scores on. Some speech collections are shared for research only, and their licence does not allow a commercial company to publish results from them. We do not publish those numbers anywhere, even when they flatter us. The data we do publish ourselves, like the script above and any results table we put on this site, is free for anyone to reuse under the Creative Commons Attribution 4.0 licence, so long as you credit DanDin and link back to the page you took it from.

Who pays for this, and who can influence it?

You should ask, so here is the answer in full. No company we compare pays us anything, gives us a free licence, or sees a page before it goes up. There are no affiliate links and no paid links on this site, so no page earns us a commission when you click through to another company. Nobody outside DanDin can add a tool to a comparison or change a line in one.

What we do gain is obvious and worth saying out loud: we sell a dictation app, and these pages exist partly so people find it. That is the bias to watch for, which is why the method, the script and the arithmetic are all on this page for you to check, and why every comparison has to name the people who should stay with the other app.

How do we handle our own product?

DanDin is our product. We apply the same method to DanDin and to every app we compare, and when we test, DanDin will read the same script on the same machine on the same day as everything else. Every comparison we publish sits on one page, all comparisons, and each one has to answer one question in plain words: who should stay with the other app? If we cannot answer that honestly, the comparison does not get published.

We are also not allowed to hide behind a number we do not have. We do not publish an accuracy figure for DanDin, or for anyone else, and we do not claim DanDin understands a list of dialects, because no such test has been run and published. What we do say is how it is built: for Arabic as people really speak it, with a step that is instructed never to turn a dialect word into formal Arabic.

How is AI used in writing these pages?

We use AI tools to help write these pages, and we would rather say so than hide it. Before a page goes live, every claim about DanDin is checked against the version people can download today, and every fact about another company against that company's own page, with the date. Numbers, prices, requirements and limits are read out of the code and the company pages they come from, not from memory. If a claim cannot be checked, it is cut, not softened.

A person also reads every page, but not always before it goes live. Each page keeps a list of the reads it still owes, and a page stays on that list until a person has read it. When a read finds a mistake, we fix the page and change its date.

Which pages have we read, and how?

Every page we read is listed at the bottom of this page with the date we read it. We quote from the raw page, the same text a browser gets, not from a summary of it, because a summary can drop a footnote. Quotes are copied, never retyped from memory. Where a company's own text uses a long dash, we print a hyphen, because this site uses none; nothing else in a quote is changed.

We re-read a company page when we touch a comparison, and the date on the page changes only then. A date on this site is never moved to look fresh.

We are also not promising you a date for the hands-on tests. A promised date becomes a lie the day it slips, and you would have no way of knowing. What we promise instead is that the words "not tested yet" stay on a page until the test has actually run, and that the test date appears on the page the day it does.

What has changed on our pages?

  • 8 October 2026. We read again every page we quote from the companies we compare. Wispr Flow had rewritten its page on languages, so we removed our old quotes from it and now quote only what it says today. DanDin itself changed too: since version 1.3.0 its default mode writes Arabic and English in one dictation, so we corrected the pages that said you had to pick one language first.

How do I tell you about a mistake?

Email support@dandin.com with the address of the page and what is wrong on it. If you are right, we change the page, change the review date on it, and say what we changed. If a whole page turns out to be wrong, we take it down rather than patch it quietly.

Common questions

Do you test competing apps yourselves?

Not yet. Every fact about another app on this site comes from that company's own pages, with the link and the date we read it. No page on this site says we tested another app, and the table on this page marks each one "not tested yet" until we run the script on it.

How can a company that makes a dictation app compare fairly?

By saying it owns one of the apps, by using the same script, the same machine and the same day for every app including its own, by publishing the method before the results, and by naming on every comparison the people who should stay with the other app. You can also check our numbers, because the script and the way we count errors are both on this page.

What is word error rate?

Count the words an app changed, dropped or added, then divide by the number of words in the correct text. Ten words with one wrong word and one missing word is two errors out of ten, which is 20 percent. Lower is better.

Why not publish results from synthetic voices?

A computer voice is not a person. It has no dialect, no accent, no hesitation and no background noise, so a score on it would tell you nothing about your own speech. We use one synthetic voice only to show what DanDin types, and we label it every time.

Do the companies you compare pay you anything?

No. No money, no free licences, no affiliate or paid links anywhere on this site. Nobody can buy a place in a comparison or a better line in one, and nobody sees a page before it is published.

How do I report a mistake on a page?

Email support@dandin.com with the page address and what is wrong. We fix the page, change the date on it, and say what changed.

Sources

  1. Google: Creating helpful, reliable, people-first content, read on
  2. Microsoft: Use voice typing to talk instead of type on your PC, read on
  3. Microsoft: Set up Voice access, read on
  4. Microsoft: Dictate your documents in Word, read on
  5. Apple: Dictate messages and documents on Mac, read on
  6. Apple: macOS feature availability, read on
  7. Apple: iOS and iPadOS feature availability, read on
  8. Google: Type with your voice in Google Docs, read on
  9. Google: Type with your voice, Gboard Help, read on
  10. Wispr Flow: Use Flow with multiple languages, read on
  11. Wispr Flow: Supported devices and system requirements, read on
  12. Wispr Flow: Pricing, read on
  13. Superwhisper: home page, read on
  14. Superwhisper: Windows, read on
  15. Superwhisper: Download, read on
  16. Nuance: Dragon Professional, read on
  17. Nuance: Dragon Legal, read on
  18. FreeFlow, the open source Mac dictation app (repository and readme), read on

Keep reading

Try DanDin today

Start free. No card.