Research · October 2026
Audicite transcription benchmark
We measured word error rate, speaker attribution and speed on 41 recordings of multi-speaker dialogue (33 minutes), clean and with background noise. Audicite made 1.38% word errors overall and put 97.3% of words under the right speaker, transcribing about 30x real time.
Updated by Audicite
Results
| Condition | System | Word error rate | Right speaker | Speed | Files |
|---|---|---|---|---|---|
| Clean | Audicite (upload pipeline) | 0.55% | 94.8% | 30x real time | 10 |
| Clean | Whisper base.en (open source, local) | 3.38% | n/a | 23x real time | 10 |
| 20 dB noise | Audicite (upload pipeline) | 0.36% | 97.6% | 31x real time | 10 |
| 20 dB noise | Whisper base.en (open source, local) | 3.20% | n/a | 25x real time | 10 |
| 10 dB noise | Audicite (upload pipeline) | 0.56% | 99.5% | 32x real time | 10 |
| 10 dB noise | Whisper base.en (open source, local) | 4.19% | n/a | 24x real time | 10 |
| 5 dB noise | Audicite (upload pipeline) | 0.83% | 96.8% | 32x real time | 10 |
| 5 dB noise | Whisper base.en (open source, local) | 9.62% | n/a | 25x real time | 10 |
| Fast overlapping speech | Audicite (upload pipeline) | 10.44% | 99.2% | 20x real time | 1 |
| Fast overlapping speech | Whisper base.en (open source, local) | 11.06% | n/a | 15x real time | 1 |
| All | Audicite (upload pipeline) | 1.38% | 97.3% | 30x real time | 41 |
| All | Whisper base.en (open source, local) | 5.59% | n/a | 23x real time | 41 |
Lower word error rate is better. “Right speaker” is the share of recognised words shown under the person who said them; Whisper does not separate speakers, so it has none. Averages are weighted by recording length.
Method
| Corpus | 11 scripted dialogues of 2 to 4 speakers (hearings, interviews, panels, a town hall, a council meeting, a debate and a fast walk-and-talk), 10.2 minutes; 10 of them also with background noise at 20, 10 and 5 dB signal-to-noise ratio |
|---|---|
| Speech | Synthetic voices with exact word-level gold transcripts and speaker labels |
| Word error rate | Word-level edit distance divided by reference words, after lowercasing, removing punctuation and writing numbers as digits on both sides |
| Speaker attribution | Share of recognised words shown under the right speaker, after the best one-to-one mapping of labels by overlap |
| Speed | Recording length divided by processing time, including upload |
| Systems | Audicite’s production upload pipeline; open-source Whisper base.en through whisper.cpp on an Apple silicon laptop |
| Date | 2026-10-06 |
Limits
- Synthetic speech is cleaner and more regular than real people: real recordings will score worse, for every system.
- English only, and a small corpus. Treat differences under one percentage point as noise.
- Live transcription uses a different engine and is not measured here.
- Commercial products other than Audicite were not tested: their accuracy cannot be measured without accounts and permission to benchmark them. We will add systems we can test fairly.
- Whisper base is a small model; larger Whisper models are more accurate and slower.
Reproduce it
The scoring code (word error rate, number normalisation, speaker mapping) and the corpus scripts are part of Audicite’s evaluation tools; the full per-file results are below. To request the corpus for your own testing, email hello@audicite.com.
Per-file results (82)
| File | System | Condition | WER | Speaker |
|---|---|---|---|---|
| d01-senate-debate | audicite | Clean | 0.00% | 99.4% |
| d01-senate-debate_snr10 | audicite | 10 dB noise | 0.00% | 99.4% |
| d01-senate-debate_snr20 | audicite | 20 dB noise | 0.00% | 99.4% |
| d01-senate-debate_snr5 | audicite | 5 dB noise | 0.00% | 99.4% |
| d02-budget-hearing | audicite | Clean | 3.51% | 100.0% |
| d02-budget-hearing_snr10 | audicite | 10 dB noise | 0.88% | 100.0% |
| d02-budget-hearing_snr20 | audicite | 20 dB noise | 1.75% | 100.0% |
| d02-budget-hearing_snr5 | audicite | 5 dB noise | 0.88% | 100.0% |
| d03-press-briefing | audicite | Clean | 1.45% | 99.3% |
| d03-press-briefing_snr10 | audicite | 10 dB noise | 1.45% | 99.3% |
| d03-press-briefing_snr20 | audicite | 20 dB noise | 1.45% | 99.3% |
| d03-press-briefing_snr5 | audicite | 5 dB noise | 1.45% | 99.3% |
| d04-election-interview | audicite | Clean | 0.00% | 100.0% |
| d04-election-interview_snr10 | audicite | 10 dB noise | 1.67% | 100.0% |
| d04-election-interview_snr20 | audicite | 20 dB noise | 0.00% | 100.0% |
| d04-election-interview_snr5 | audicite | 5 dB noise | 1.67% | 100.0% |
| d05-parliament-panel | audicite | Clean | 0.00% | 79.3% |
| d05-parliament-panel_snr10 | audicite | 10 dB noise | 0.00% | 99.1% |
| d05-parliament-panel_snr20 | audicite | 20 dB noise | 0.00% | 99.1% |
| d05-parliament-panel_snr5 | audicite | 5 dB noise | 0.00% | 99.1% |
| d06-town-hall | audicite | Clean | 0.00% | 100.0% |
| d06-town-hall_snr10 | audicite | 10 dB noise | 0.93% | 100.0% |
| d06-town-hall_snr20 | audicite | 20 dB noise | 0.00% | 100.0% |
| d06-town-hall_snr5 | audicite | 5 dB noise | 0.93% | 100.0% |
| d07-trade-roundtable | audicite | Clean | 0.00% | 88.6% |
| d07-trade-roundtable_snr10 | audicite | 10 dB noise | 0.00% | 98.9% |
| d07-trade-roundtable_snr20 | audicite | 20 dB noise | 0.00% | 98.8% |
| d07-trade-roundtable_snr5 | audicite | 5 dB noise | 0.00% | 88.6% |
| d08-city-council | audicite | Clean | 0.00% | 100.0% |
| d08-city-council_snr10 | audicite | 10 dB noise | 0.00% | 100.0% |
| d08-city-council_snr20 | audicite | 20 dB noise | 0.00% | 100.0% |
| d08-city-council_snr5 | audicite | 5 dB noise | 0.00% | 100.0% |
| d09-legislative-briefing | audicite | Clean | 0.00% | 100.0% |
| d09-legislative-briefing_snr10 | audicite | 10 dB noise | 0.00% | 100.0% |
| d09-legislative-briefing_snr20 | audicite | 20 dB noise | 0.00% | 100.0% |
| d09-legislative-briefing_snr5 | audicite | 5 dB noise | 0.00% | 100.0% |
| d10-debate-rebuttal | audicite | Clean | 0.00% | 77.2% |
| d10-debate-rebuttal_snr10 | audicite | 10 dB noise | 0.00% | 98.0% |
| d10-debate-rebuttal_snr20 | audicite | 20 dB noise | 0.00% | 77.2% |
| d10-debate-rebuttal_snr5 | audicite | 5 dB noise | 2.97% | 77.2% |
| d10-walk-and-talk | audicite | Fast overlapping speech | 10.44% | 99.2% |
| d01-senate-debate | whisper | Clean | 1.12% | n/a |
| d01-senate-debate_snr10 | whisper | 10 dB noise | 1.69% | n/a |
| d01-senate-debate_snr20 | whisper | 20 dB noise | 1.12% | n/a |
| d01-senate-debate_snr5 | whisper | 5 dB noise | 1.69% | n/a |
| d02-budget-hearing | whisper | Clean | 7.02% | n/a |
| d02-budget-hearing_snr10 | whisper | 10 dB noise | 5.26% | n/a |
| d02-budget-hearing_snr20 | whisper | 20 dB noise | 5.26% | n/a |
| d02-budget-hearing_snr5 | whisper | 5 dB noise | 14.04% | n/a |
| d03-press-briefing | whisper | Clean | 5.80% | n/a |
| d03-press-briefing_snr10 | whisper | 10 dB noise | 10.14% | n/a |
| d03-press-briefing_snr20 | whisper | 20 dB noise | 3.62% | n/a |
| d03-press-briefing_snr5 | whisper | 5 dB noise | 19.57% | n/a |
| d04-election-interview | whisper | Clean | 5.83% | n/a |
| d04-election-interview_snr10 | whisper | 10 dB noise | 8.33% | n/a |
| d04-election-interview_snr20 | whisper | 20 dB noise | 9.17% | n/a |
| d04-election-interview_snr5 | whisper | 5 dB noise | 10.83% | n/a |
| d05-parliament-panel | whisper | Clean | 0.00% | n/a |
| d05-parliament-panel_snr10 | whisper | 10 dB noise | 0.93% | n/a |
| d05-parliament-panel_snr20 | whisper | 20 dB noise | 0.00% | n/a |
| d05-parliament-panel_snr5 | whisper | 5 dB noise | 2.78% | n/a |
| d06-town-hall | whisper | Clean | 2.80% | n/a |
| d06-town-hall_snr10 | whisper | 10 dB noise | 2.80% | n/a |
| d06-town-hall_snr20 | whisper | 20 dB noise | 0.00% | n/a |
| d06-town-hall_snr5 | whisper | 5 dB noise | 5.61% | n/a |
| d07-trade-roundtable | whisper | Clean | 0.00% | n/a |
| d07-trade-roundtable_snr10 | whisper | 10 dB noise | 1.18% | n/a |
| d07-trade-roundtable_snr20 | whisper | 20 dB noise | 0.00% | n/a |
| d07-trade-roundtable_snr5 | whisper | 5 dB noise | 3.53% | n/a |
| d08-city-council | whisper | Clean | 0.97% | n/a |
| d08-city-council_snr10 | whisper | 10 dB noise | 0.00% | n/a |
| d08-city-council_snr20 | whisper | 20 dB noise | 1.94% | n/a |
| d08-city-council_snr5 | whisper | 5 dB noise | 0.97% | n/a |
| d09-legislative-briefing | whisper | Clean | 1.61% | n/a |
| d09-legislative-briefing_snr10 | whisper | 10 dB noise | 0.00% | n/a |
| d09-legislative-briefing_snr20 | whisper | 20 dB noise | 0.00% | n/a |
| d09-legislative-briefing_snr5 | whisper | 5 dB noise | 0.00% | n/a |
| d10-debate-rebuttal | whisper | Clean | 6.93% | n/a |
| d10-debate-rebuttal_snr10 | whisper | 10 dB noise | 7.92% | n/a |
| d10-debate-rebuttal_snr20 | whisper | 20 dB noise | 8.91% | n/a |
| d10-debate-rebuttal_snr5 | whisper | 5 dB noise | 33.66% | n/a |
| d10-walk-and-talk | whisper | Fast overlapping speech | 11.06% | n/a |
Questions
How accurate is Audicite?
On this benchmark Audicite’s word error rate was 0.55% on clean dialogues, 0.83% with heavy noise (5 dB), and 10.4% on fast overlapping speech; 97.3% of words were attributed to the right speaker overall.
Which transcription tool is most accurate?
It depends on the audio. Measure on recordings like yours: this page publishes the method, and the corpus and scoring code can be reused. On this corpus Audicite’s error rate was about a quarter of open-source Whisper base’s.
More
Published by Audicite. We make one of the systems tested; see the method and limits above, and about us.