Skip to main content
Benchmarked · August 2026

Measured against other systems

We ran our self-hosted engine and every Arabic-focused competitor we could reach through a public API over the same real, bilingual meeting audio, and scored them all the same way. Here is the whole result — including the Arabic column, where three of the four competing configurations beat us.

Results

35.14

Overall word error rate

The lowest overall error rate in this comparison — but we don't claim a win. Two of the four paired intervals span zero, so against the leading vendors we report parity rather than a ranking.

+16.3+34.7points

English words inside Arabic speech

More accurate on the English words spoken mid-Arabic-sentence than each of the four competing configurations, one by one — paired intervals at 95% confidence, all excluding zero. This is the one axis where we claim the lead.

>3×

Arabic–English switch points kept

When someone switches between Arabic and English mid-sentence, more of those moments survive into the transcript — over three times as many as any of the three Speechmatics configurations kept. The error counted is a script collapse: an English word written in Arabic letters.

Every figure here describes the self-hosted configuration, running on your own hardware. We ran this comparison ourselves on 80 code-switched utterances of real UN ESCWA meeting audio. It is not a measurement of our hosted tiers.

One dataset: real UN ESCWA meeting audio — 80 utterances, every one code-switched.

Swipe the table to see every column.

Benchmark comparison of speech recognition systems on identical audio
SystemOverall WERLower is betterArabic wordsLower is betterEnglish wordsLower is betterSwitch points keptHigher is better
Khulasa — self-hosted35.1437.1629.0839.0%
Speechmaticsstandard Arabic/English37.2934.6945.4011.0%
Speechmatics melia-1code-switch model38.6030.2663.8112.1%
Speechmaticsenhanced Arabic/English40.1832.3763.187.4%
A leading Arabic cloud vendorcode-switch model43.8540.3954.6025.0%
Khulasa — hosted tiersMeasurement pending

We do not claim to win this column. Two of the four paired intervals span zero, so at this sample size we report parity with the leading vendors rather than a ranking. We do not win the Arabic column either — all three Speechmatics configurations are ahead of us there. The English column is the one where every interval excludes zero.

Our hosted tiers use managed cloud models and have not been measured on this audio yet. No figure is shown for them until they are.

How to read the intervals

Every comparison here is paired: each system transcribed the same utterances, and the difference between two systems is resampled 2,000 times at 95% confidence. Where that interval spans zero, the two systems are not distinguishable at this sample size, and we say so rather than naming a winner. That is why the overall column is reported as parity, while the English column — where every interval excludes zero — is the one axis where we claim the lead.

English words inside Arabic speech
Paired intervals at 95% confidence, all excluding zero — most conservative lower bound+8.8
Arabic–English switch points kept
39.0%vs7.4%12.1%(Speechmatics)

By dialect

We make no per-dialect claim yet. Each dialect below shows what has been measured and published for it, and nothing more.

Gulf

  • Saudi — NajdiToo few clips for this data to resolve a result
  • Saudi — HijaziToo few clips for this data to resolve a result
  • EmiratiMeasured, not yet published
  • Gulf — otherToo few clips for this data to resolve a result

Levant

  • LevantineNot yet measured

Iraq

  • IraqiNot yet measured

Nile Valley

  • EgyptianNot covered by the current data
  • SudaneseNot yet measured

Maghreb

  • MoroccanNot yet measured
  • AlgerianNot yet measured
  • TunisianNot yet measured

Other

  • YemeniNot yet measured
  • Modern Standard ArabicNot yet measured

Method

Why the English column decides an Arabic meeting

An Arabic meeting is not an Arabic-only meeting. The product names, the tools, the metrics, the deadlines — the words an action item is actually about — arrive in English, mid-sentence, in Latin script.

Overall word error rate is dominated by the Arabic tokens, because in this audio there are more of them. A system can therefore score well overall while writing the English words in Arabic letters: the aggregate barely moves, and the English word is gone from the transcript. Nothing downstream can recover a word that was never written down.

So read the table as a trade, not a sweep. All three Speechmatics configurations transcribe the Arabic words more accurately than this one does; on the English words this one is far ahead of every system here; the overall rate comes out level. In a bilingual meeting that is the trade we would take — and it is the axis we measured and published.

Across the four competing configurations the collapse count runs from 38 to 164, the highest being Speechmatics melia-1; this configuration collapsed 53, which is not the lowest figure in the comparison.

The model is open — how we run it is ours

The speech model is Whisper large-v3 — from OpenAI, MIT licence, open weights. We did not train it, and we lead with that. What is ours is everything around it: the model runs at the published reference configuration its authors describe rather than at the library default, language is decided once per recording instead of once per turn, and the deployment is on-premise, so the audio stays on your own hardware. Then we measured it on bilingual meetings and published what came back.

Whisper large-v3OpenAIMIT

How this was measured

We ran this comparison ourselves. Every system transcribed identical audio and was scored by the same normalizer. We report corpus-level word error rate with a paired per-utterance bootstrap — 2,000 draws at 95% confidence — and where an interval spans zero we call the systems not distinguishable rather than naming a winner. The audio is 80 utterances of real UN ESCWA meeting audio, every one code-switched, selected by a hash of the utterance identifier rather than chosen by us. Every figure describes the self-hosted configuration of Khulasa, the one that runs on your own hardware, on-premise. Measured August 2026.

Corpora

Used under their published licences. Only the ESCWA audio carries the figures in the table; the others are named because the limits below draw on them.

Text normalization

Every system's output and every reference transcript were scored by the same normalizer.

What we do not claim

  • These are public research corpora, not customer meetings. A real meeting is harder than all of them, and no customer audio has ever been measured.
  • One corpus of 80 utterances is a small sample, which is why we publish intervals rather than rankings.
  • We make no per-dialect claim — there are too few clips per dialect for this data to resolve one, and Egyptian Arabic is not covered at all.
  • On dialectal Arabic that is not meeting audio, the leading vendors are ahead of this configuration — by around 14 points on Saudi broadcast. We chose this configuration for bilingual meetings, and that is the trade we made.
  • Word error rate is not summary quality. A system can score well here and still produce a fluent, confident, wrong sentence, and nothing on this page measures the summary.
  • Every figure here describes the self-hosted, on-premise configuration — the one that runs on your own hardware. This is not a measurement of our hosted tiers.
  • On English-only meeting audio the leading vendor is ahead of this configuration — by around 4 to 6 points. The advantage we report is specific to English words spoken inside Arabic sentences.
  • Intella and Notah are not in this table. Neither exposes a public API we could run this audio through, so neither could be measured here.

Scope and trademarks

  • Not every system here is named. Where a name is not needed to identify the configuration measured, we describe it instead; the rules on comparative naming differ across the markets this page is read in. Every system was run, scored and reported the same way, and the label changes no figure.
  • Each competing system was run through its vendor's public API, in the configuration shown in the table, on the audio described above, in August 2026. Our own row is the self-hosted engine at its shipped configuration.
  • Every system in this table is a bilingual Arabic-and-English configuration; no Arabic-only result is reported here. Where a vendor offers more than one such configuration we ran more than one: three Speechmatics configurations were run on this audio and all three are shown, and for the vendor described rather than named we ran every model its API exposes and report its best result.
  • These figures describe those configurations, on that audio, on that date. Vendors update their models, and the same run repeated later may not reproduce them.
  • Each corpus above is used under the licence printed beside it, and several are restricted to non-commercial research. What we publish are measurements computed from that audio; we redistribute no audio and no reference transcript, in whole or in part.
  • Product and company names are the trademarks of their respective owners. Khulasa is not affiliated with, sponsored by or endorsed by any of them, and no vendor reviewed or approved these results.
  • If a vendor believes a configuration shown here is not their best available, tell us and we will re-run the comparison.

Changelog

Last updated August 2026.

Benchmark changelog

Run it on your own hardware

Every figure on this page describes the self-hosted configuration. Size an on-premise or private-cloud deployment with us.

Cookies on Khulasa

We use strictly necessary cookies to run the site. With your permission we also use Google Analytics to understand how the site is used. You can change your choice at any time from "Cookie settings" in the footer. Privacy Policy