Skip to content
Audicite

Research · October 2026

Audicite transcription benchmark

We measured word error rate, speaker attribution and speed on 41 recordings of multi-speaker dialogue (33 minutes), clean and with background noise. Audicite made 1.38% word errors overall and put 97.3% of words under the right speaker, transcribing about 30x real time.

Updated by Audicite

Results

Benchmark results by condition
ConditionSystemWord error rateRight speakerSpeedFiles
CleanAudicite (upload pipeline)0.55%94.8%30x real time10
CleanWhisper base.en (open source, local)3.38%n/a23x real time10
20 dB noiseAudicite (upload pipeline)0.36%97.6%31x real time10
20 dB noiseWhisper base.en (open source, local)3.20%n/a25x real time10
10 dB noiseAudicite (upload pipeline)0.56%99.5%32x real time10
10 dB noiseWhisper base.en (open source, local)4.19%n/a24x real time10
5 dB noiseAudicite (upload pipeline)0.83%96.8%32x real time10
5 dB noiseWhisper base.en (open source, local)9.62%n/a25x real time10
Fast overlapping speechAudicite (upload pipeline)10.44%99.2%20x real time1
Fast overlapping speechWhisper base.en (open source, local)11.06%n/a15x real time1
AllAudicite (upload pipeline)1.38%97.3%30x real time41
AllWhisper base.en (open source, local)5.59%n/a23x real time41

Lower word error rate is better. “Right speaker” is the share of recognised words shown under the person who said them; Whisper does not separate speakers, so it has none. Averages are weighted by recording length.

Method

Benchmark method
Corpus11 scripted dialogues of 2 to 4 speakers (hearings, interviews, panels, a town hall, a council meeting, a debate and a fast walk-and-talk), 10.2 minutes; 10 of them also with background noise at 20, 10 and 5 dB signal-to-noise ratio
SpeechSynthetic voices with exact word-level gold transcripts and speaker labels
Word error rateWord-level edit distance divided by reference words, after lowercasing, removing punctuation and writing numbers as digits on both sides
Speaker attributionShare of recognised words shown under the right speaker, after the best one-to-one mapping of labels by overlap
SpeedRecording length divided by processing time, including upload
SystemsAudicite’s production upload pipeline; open-source Whisper base.en through whisper.cpp on an Apple silicon laptop
Date2026-10-06

Limits

  • Synthetic speech is cleaner and more regular than real people: real recordings will score worse, for every system.
  • English only, and a small corpus. Treat differences under one percentage point as noise.
  • Live transcription uses a different engine and is not measured here.
  • Commercial products other than Audicite were not tested: their accuracy cannot be measured without accounts and permission to benchmark them. We will add systems we can test fairly.
  • Whisper base is a small model; larger Whisper models are more accurate and slower.

Reproduce it

The scoring code (word error rate, number normalisation, speaker mapping) and the corpus scripts are part of Audicite’s evaluation tools; the full per-file results are below. To request the corpus for your own testing, email hello@audicite.com.

Per-file results (82)
FileSystemConditionWERSpeaker
d01-senate-debateaudiciteClean0.00%99.4%
d01-senate-debate_snr10audicite10 dB noise0.00%99.4%
d01-senate-debate_snr20audicite20 dB noise0.00%99.4%
d01-senate-debate_snr5audicite5 dB noise0.00%99.4%
d02-budget-hearingaudiciteClean3.51%100.0%
d02-budget-hearing_snr10audicite10 dB noise0.88%100.0%
d02-budget-hearing_snr20audicite20 dB noise1.75%100.0%
d02-budget-hearing_snr5audicite5 dB noise0.88%100.0%
d03-press-briefingaudiciteClean1.45%99.3%
d03-press-briefing_snr10audicite10 dB noise1.45%99.3%
d03-press-briefing_snr20audicite20 dB noise1.45%99.3%
d03-press-briefing_snr5audicite5 dB noise1.45%99.3%
d04-election-interviewaudiciteClean0.00%100.0%
d04-election-interview_snr10audicite10 dB noise1.67%100.0%
d04-election-interview_snr20audicite20 dB noise0.00%100.0%
d04-election-interview_snr5audicite5 dB noise1.67%100.0%
d05-parliament-panelaudiciteClean0.00%79.3%
d05-parliament-panel_snr10audicite10 dB noise0.00%99.1%
d05-parliament-panel_snr20audicite20 dB noise0.00%99.1%
d05-parliament-panel_snr5audicite5 dB noise0.00%99.1%
d06-town-hallaudiciteClean0.00%100.0%
d06-town-hall_snr10audicite10 dB noise0.93%100.0%
d06-town-hall_snr20audicite20 dB noise0.00%100.0%
d06-town-hall_snr5audicite5 dB noise0.93%100.0%
d07-trade-roundtableaudiciteClean0.00%88.6%
d07-trade-roundtable_snr10audicite10 dB noise0.00%98.9%
d07-trade-roundtable_snr20audicite20 dB noise0.00%98.8%
d07-trade-roundtable_snr5audicite5 dB noise0.00%88.6%
d08-city-councilaudiciteClean0.00%100.0%
d08-city-council_snr10audicite10 dB noise0.00%100.0%
d08-city-council_snr20audicite20 dB noise0.00%100.0%
d08-city-council_snr5audicite5 dB noise0.00%100.0%
d09-legislative-briefingaudiciteClean0.00%100.0%
d09-legislative-briefing_snr10audicite10 dB noise0.00%100.0%
d09-legislative-briefing_snr20audicite20 dB noise0.00%100.0%
d09-legislative-briefing_snr5audicite5 dB noise0.00%100.0%
d10-debate-rebuttalaudiciteClean0.00%77.2%
d10-debate-rebuttal_snr10audicite10 dB noise0.00%98.0%
d10-debate-rebuttal_snr20audicite20 dB noise0.00%77.2%
d10-debate-rebuttal_snr5audicite5 dB noise2.97%77.2%
d10-walk-and-talkaudiciteFast overlapping speech10.44%99.2%
d01-senate-debatewhisperClean1.12%n/a
d01-senate-debate_snr10whisper10 dB noise1.69%n/a
d01-senate-debate_snr20whisper20 dB noise1.12%n/a
d01-senate-debate_snr5whisper5 dB noise1.69%n/a
d02-budget-hearingwhisperClean7.02%n/a
d02-budget-hearing_snr10whisper10 dB noise5.26%n/a
d02-budget-hearing_snr20whisper20 dB noise5.26%n/a
d02-budget-hearing_snr5whisper5 dB noise14.04%n/a
d03-press-briefingwhisperClean5.80%n/a
d03-press-briefing_snr10whisper10 dB noise10.14%n/a
d03-press-briefing_snr20whisper20 dB noise3.62%n/a
d03-press-briefing_snr5whisper5 dB noise19.57%n/a
d04-election-interviewwhisperClean5.83%n/a
d04-election-interview_snr10whisper10 dB noise8.33%n/a
d04-election-interview_snr20whisper20 dB noise9.17%n/a
d04-election-interview_snr5whisper5 dB noise10.83%n/a
d05-parliament-panelwhisperClean0.00%n/a
d05-parliament-panel_snr10whisper10 dB noise0.93%n/a
d05-parliament-panel_snr20whisper20 dB noise0.00%n/a
d05-parliament-panel_snr5whisper5 dB noise2.78%n/a
d06-town-hallwhisperClean2.80%n/a
d06-town-hall_snr10whisper10 dB noise2.80%n/a
d06-town-hall_snr20whisper20 dB noise0.00%n/a
d06-town-hall_snr5whisper5 dB noise5.61%n/a
d07-trade-roundtablewhisperClean0.00%n/a
d07-trade-roundtable_snr10whisper10 dB noise1.18%n/a
d07-trade-roundtable_snr20whisper20 dB noise0.00%n/a
d07-trade-roundtable_snr5whisper5 dB noise3.53%n/a
d08-city-councilwhisperClean0.97%n/a
d08-city-council_snr10whisper10 dB noise0.00%n/a
d08-city-council_snr20whisper20 dB noise1.94%n/a
d08-city-council_snr5whisper5 dB noise0.97%n/a
d09-legislative-briefingwhisperClean1.61%n/a
d09-legislative-briefing_snr10whisper10 dB noise0.00%n/a
d09-legislative-briefing_snr20whisper20 dB noise0.00%n/a
d09-legislative-briefing_snr5whisper5 dB noise0.00%n/a
d10-debate-rebuttalwhisperClean6.93%n/a
d10-debate-rebuttal_snr10whisper10 dB noise7.92%n/a
d10-debate-rebuttal_snr20whisper20 dB noise8.91%n/a
d10-debate-rebuttal_snr5whisper5 dB noise33.66%n/a
d10-walk-and-talkwhisperFast overlapping speech11.06%n/a

Questions

How accurate is Audicite?

On this benchmark Audicite’s word error rate was 0.55% on clean dialogues, 0.83% with heavy noise (5 dB), and 10.4% on fast overlapping speech; 97.3% of words were attributed to the right speaker overall.

Which transcription tool is most accurate?

It depends on the audio. Measure on recordings like yours: this page publishes the method, and the corpus and scoring code can be reused. On this corpus Audicite’s error rate was about a quarter of open-source Whisper base’s.

More

Published by Audicite. We make one of the systems tested; see the method and limits above, and about us.