Vendor benchmarks have a credibility problem: everyone publishes the numbers
they win. So when we measured Audioscrape against Deepgram nova-3 and
AssemblyAI universal-3-5-pro, we set one rule: publish the whole grid,
including the columns we lose, and every claim of ours that died along the
way.
Everything ran on 29 public recordings across three domains (AMI meetings,
Earnings-21 business calls, VoxConverse in-the-wild clips), scored with
open-source scorers (pyannote.metrics, jiwer) applied identically to every
system. Our numbers come from the production pipeline, the one customers use.
Full tables on the benchmarks page.
Claim one: died in a robustness sweep
Our first diarization comparison showed a 2–3× better diarization error rate than Deepgram. It looked great, and it was wrong: utterance-level DER measures transcript formatting, not diarization. One vendor merges speech so aggressively that its utterances claimed 15.9 minutes of speech in a 13.1-minute meeting.
Re-scored from word timings (the common denominator all three systems report) with the one free parameter swept across its range, our lead over Deepgram reversed at one end of the sweep. A lead that only exists at one parameter setting is not a lead. We dropped the claim.
Claim two: died when the corpus grew
On our first fixture we beat both vendors on word error rate, and for two days the sales line was “we’re ahead on transcription too.” Then we expanded from one meeting to 29 recordings: Deepgram takes meetings-WER, AssemblyAI takes business-call WER, both by under a point, and per-recording winners flip constantly. One recording is an anecdote, not a benchmark.
So here is the honest sentence about raw transcription: at the top of this field, everyone is within a point or two, and nobody wins everywhere. That is exactly why raw transcription is the wrong axis to buy on.
Claim three: died to two German clips
“The best in-the-wild diarization of the three” was true right up until we looked at why one column was so bad.
Two of the ten VoxConverse clips are German. The Deepgram model we called is English-only, so it returned one and five words for them. Our scorer treated those empty transcripts as ordinary hypotheses and charged Deepgram a ~100% miss — for a language gap, not for diarization. Nothing errored. The number just looked like a win.
Removing those two clips moved Deepgram’s in-the-wild DER from 24.1% to 5.2%, and turned our win into a loss by more than 2x. Deepgram is simply better than us on this kind of audio, and the table now says so.
We found it because a fourth system’s results made us compare per-recording word counts across systems for the first time. That check is now part of the harness: if any system returns under 30% of what the others returned on the same recording, the run fails instead of scoring. An empty answer scores badly rather than loudly, which is exactly how a benchmark misleads the people who built it.
The same correction cut the other way too, and we will take it with the same face: excluding those clips, our diarization now beats AssemblyAI in all nine domain-by-setting cells rather than eight of nine.
What survived the sweeps
- Better overall diarization than AssemblyAI on meetings and on business calls, at every setting we swept (0 to 2.0s). Not in-the-wild: we lead there across the range we publish, and AssemblyAI passes us once the merge rule opens past 1.0s. We found that by sweeping further than we print, which is the only reason we can tell you about it.
- A wrong-mouth rate ~2x lower than Deepgram on meetings and on business calls: the confusion component of DER, the one that means “these words were attributed to the wrong person.” Deepgram won zero of the eight meetings on it. For a compliance review or an investigation, that is the number that matters. In the wild Deepgram beats us on every column, and that is on the page too.
And the one that isn’t a number contest at all: identity persistence. Label a voice once in your workspace, and later recordings of the same group arrive already named: 87% of speech at 88.9% precision in our enroll-once test, stable across a 2× threshold range. Per-file APIs score 0% here by construction: they return “Speaker 0” with nothing connecting it to last month’s meeting. It’s not a number they lose; it’s a capability the per-file architecture doesn’t have. And it’s “label once, done”: the learning curve on four-voice casts is flat, so we say that instead of the marketing line the data doesn’t support.
Why publish losing columns?
Because the test you should run on every vendor benchmark, including ours, is: what did they measure and decide not to show you? We’d rather you ask it of us first. Vendor models and run dates are printed next to every table; vendors ship silent model updates, so any number older than a quarter deserves a re-run.
If the benchmark that actually matters to you is your own audio: we’ll set up an evaluation workspace with your recordings and your speakers, and you score it with your own reviewers. Talk to us.