How I chose a speech-to-text engine

Six engines, 1,000+ of my own dictations, one noisy room. A dictation in the wrong language or speech that never arrived disqualified an engine before its error count was read.

Quick answer

The engine with the fewest errors did not become my default. I tested six speech-to-text engines on my own dictations in a noisy room, and the worst failures were not wrong words. Some engines dropped parts of what I said, and my old model sometimes returned a whole dictation in another language. I kept three, Soniox, ElevenLabs Scribe v2 and OpenAI GPT Transcribe, and made GPT Transcribe, with its flat per-second bill and identical answers on repeats, the default.

Published
Reading time
19 min
Engines
6

I tested six speech-to-text engines on 1,000+ of my own Ukrainian dictations, about 10 hours of speech. The one with the fewest errors did not become my default.

The model Shepit had been running on faced Soniox, ElevenLabs Scribe v2 and OpenAI GPT Transcribe. Deepgram Nova-2 and Nova-3 and Speechmatics Linden 1 joined them. Three stayed, two were cut, and the old model dropped to last resort. The error rate ranked the survivors, but it did not choose them.

If you are choosing a speech-to-text API for a language other than English, this test is for you, because the error rate hides the failures that matter.

Why I looked

Shepit has one job: hold the key, speak, release, and the words land in the field. The engine under that job was OpenAI gpt-4o-mini-transcribe, mini from here on, and mostly it did the job. Then I read my own transcripts properly, and four problems kept turning up.

It changed language on me. I dictate in Ukrainian, yet on 8 of the 76 test clips the whole dictation came back in another language, not just a word. Pinning the language in the request did not help: on the clips where mini had drifted before, 37 of 45 repeats drifted again.

It transcribed silence. Tap the hotkey by accident, say nothing, and mini turns 3 s of room noise into text: Shepit's own instruction prompt.

It never answered the same way twice. I ran 9 clips 5 times each, and not one of the 9 came back identical across its repeats. Soniox, GPT Transcribe and Deepgram gave identical text on 9 of 9.

It cut the middle out of long recordings. A 900 s file came back missing everything from 200 s to 660 s, so only 36% of the words arrived. HTTP 200, no error, no apology. The first and last lines looked fine, so a glance at the ends would never catch it.

The old model failed in ways a user cannot see or fix, not in how many words it misheard.

None of these is a wrong word. You see a wrong word and fix it, while a dictation in the wrong language, a transcript of nothing or a hole in the middle turn up later, if at all.

So I did not go into the test asking which engine mishears least. I asked which engine never does any of this, and only then, among the survivors, which one mishears least.

The scoreboard: blockers first, error count second

I wrote the rules down before I saw a single score, so I could not argue with them later. Two failures disqualify an engine outright: a whole dictation returned in the wrong language, and speech that never arrives. Both are worse than a wrong word, which at least shows up on screen. One more check applies: the engine must not invent text when I say nothing.

The rest of my brief was short:

  • No language drift and no dropped speech: critical.
  • Streaming, text that appears while I speak: very important.
  • Price: a target near $0.20 per audio hour, a hard limit of $0.30.

A blocker disqualified an engine, whatever its error rate.

76 of my hardest Ukrainian dictations; errors per 100 words after my rulings. Coloured: the cell that cut or replaced the engine.

EngineDictations in another language, of 76Clips with lost speech, of 76Silence checkErrors / 100 wordsVerdict
Soniox00Yes, one word on digital silence0.26Kept
ElevenLabs Scribe v200Yes, with sound tags off1.29Kept
OpenAI GPT Transcribe01Yes, empty1.59Kept, the default
OpenAI gpt-4o-mini-transcribe81No, returned its own prompt4.16Replaced, last resort
Deepgram06Yes, empty5.18Cut: lost speech
Speechmatics00Yes, empty8.70Cut: error count, chopped sentences

Soniox shows the file-mode request I ship; every other row is the engine's best setup: ElevenLabs Scribe v2 with Ukrainian set and sound tags removed, GPT Transcribe in file mode, mini as Shepit shipped it, Deepgram Nova-3 on finished files and Speechmatics Linden 1 streaming, all but mini with my list of 54 names and terms, of which Soniox's shipped request sends the first 50. Soniox's 0.26 is a lower bound, and its best setup, with extra context, scored 0.30 but costs $0.36 per audio hour, over my limit.

Read the table by column, not by row. With the language locked, Deepgram and Speechmatics never switched language once, and I still cut both.

Soniox made the fewest errors, and that alone did not buy it a pass: it still had to clear the silence check before I kept it. GPT Transcribe, the engine I made the default, sits only third in the error column.

How I tested

Every clip in this test is mine: real dictations from real use, not a synthetic benchmark. AI agents made the scale possible, and I kept the last word.

dictations in the corpus
1,000+
of my speech in the corpus
10 h
transcriptions
4,442+
of audio transcribed
38.6 h
of agents
dozens

One speaker, one microphone, one room; dictations of 14-22 September 2026; September 2026 model versions.

The 4,442+ is a lower bound, because refused and retried calls are not counted. It is the sum of four parts:

  • 3,543 in the two comparison rounds, 1,200 of them streaming sessions
  • 635 on the requests I ship
  • 243 in a check of my list of names and terms
  • 21 files of growing length, up to an hour each

The clips

The dictations in the corpus were recorded from 14 to 22 September 2026. I screened 1,000 of them and picked 71: the clips where mini had failed or was most likely to, and a control group.

  • 9 it had already switched language on
  • the 20 noisiest
  • 20 short ones
  • 12 with English names inside Ukrainian
  • 10 ordinary clips as the control

I added 5 long recordings, up to 6.2 minutes, for 76 test clips in all, and put 3 silence clips beside them. "Noisy" is measured, not felt: those 20 clips have the lowest signal-to-noise ratio, the gap between my voice and the room's hum.

The 76 are where all six engines met head to head. Beyond them I also checked the requests I ship, files up to an hour long and the app's own timings.

What counts as an error

The usual yardstick is word error rate (WER). Mine counts only content words, the ones that carry meaning, heard wrong, dropped or invented, and ignores small grammar words, word endings and spelling. Every rate below is errors per 100 words of the reference text, the correct transcript each engine is scored against.

The reference text

I built my first reference text from mini's own output and checked every disputed word against Soniox. On that reference, Soniox's best setup scored 0.04 errors per 100 words. Suspiciously close to perfect: my own experiment had slipped it part of the answer key. So I ruled 17 disputed clips myself and set 4 global rules, and Soniox's best setup went from 0.04 to 0.30 while mini's went from 3.22 to 4.16.

The agents and I

AI agents built and ran the pipeline, from picking the clips to scoring them. Dozens of them worked with me for about a day and a night: they researched by day, ran experiments overnight and brought verdicts in the morning. They scored 3,502 outputs from the two comparison rounds the same way and surfaced 127 disputed words.

The last word stayed mine. My 17 clip rulings and 4 global rules override the agents everywhere, and Speechmatics' 8.70 outlier got a spot check by hand.

Blocker: the wrong language

A whole dictation in another language is a broken dictation, not one wrong word. No keystroke fixes it: you delete it and speak again, if you notice.

Mini's switches were no accident of the test. 7 of its 8 landed on the 9 clips where it had already drifted in real use, so the same clips failed the same way.

whole dictations in another language
8 then 0 of 76

The old model against every kept engine, best setups, the language hint sent.

The two cut engines switched no dictation wholesale either, and the kept three hold that on the requests I ship too. What remains is one Russian-looking word form on 1 or 2 clips, scored as a single error.

The language hint carries real weight. Without its language code, ElevenLabs switched 6 of 44 clips wholesale, Croatian and Slovak among them, though I have never dictated a word of either. With the code sent, 76 of 76 came back Ukrainian. Mini is the exception: the same hint did not stop its drift.

GPT Transcribe has two modes. File mode transcribes the finished recording, and realtime mode transcribes while I speak, cutting the audio into pieces wherever the provider's voice detection hears a pause. That cutting brings the drift back: in realtime mode it switched 1 clip wholesale, while in file mode the same model switched none.

Blocker: speech that never arrives

Lost speech is the worst failure, because nobody notices it. A wrong word is on screen. A missing sentence is not.

Deepgram's best setup lost speech on 6 of 76 clips: 1 came back empty, 2 missed the start and 3 missed the end. Plain Nova-2 returned 3 clips empty. Its language never drifted, and I still cut it, on Ukrainian dictation, one speaker, one room and September 2026 model versions.

Realtime voice detection has its own way to lose a dictation. On a very quiet recording, at -71 dBFS, far below normal speech level, the provider's detection never fired. The audio was never sent on for transcription, and no error came back. The whole recording was gone.

The worst form is a hole in the middle, and mini left one in two long files. Both times it answered HTTP 200 with no signal. A hell of a success code. Shepit's own truncation check, which watches the length of mini's answer, misses such cases. If you test a speech-to-text API, check the middle of long files.

What came back from long recordings: a hole in the middle for mini, the whole file for the kept three

My files of growing length, up to 3,595 s; September 2026 model versions. Coloured: the speech that never arrived. Bold: the whole file came back.

text came backspeech missing from the textrefused0102030405060 minmini, 900 s200-660 s missingmini, 1,200 s420-1,140 s missingmini, 1,500 srefusedSoniox, 3,595 scompleteElevenLabs, 3,595 scompleteGPT Transcribe, 3,595 scomplete
text came backspeech missing from the textrefused0102030405060 minmini, 900 s200-660 s missingmini, 1,200 s420-1,140 s missingmini, 1,500 srefusedSoniox, 3,595 scompleteElevenLabs, 3,595 scompleteGPT Transcribe, 3,595 scomplete
text came backspeech missing from the textrefused01020304060 minmini, 900 s200-660 s missingmini, 1,200 s420-1,140 s missingmini, 1,500 srefusedSoniox, 3,595 scompleteElevenLabs, 3,595 scompleteGPT Transcribe, 3,595 scomplete

Mini's text begins where the recording begins and ends where it ends, so a reader who checks the first and last lines sees nothing wrong: the hole is in the middle. At 1,500 s it refused the file outright, while the three kept engines returned 3,595 s complete.

The files and what came back
Engine, fileText came back, sSpeech missing, sWeakest minute, share of wordsOutcome
mini, 900 s0-200, 660-900200-660-Hole in the middle
mini, 1,200 s0-420, 1,140-1,200420-1,140-Hole in the middle
mini, 1,500 snone--Refused
Soniox, 3,595 s0-3,595-0.96Complete
ElevenLabs, 3,595 s0-3,595-0.95Complete
GPT Transcribe, 3,595 s0-3,595-0.84Complete

The kept three passed. Soniox and ElevenLabs lost speech on 0 of 76 clips, and GPT Transcribe on 1, a cut tail.

Then I fed them files of growing length, up to an hour minus five seconds, and all three came back complete. Their weakest minute, the one where the fewest words came back, fell at a different point in the file for each engine.

Error count: quiet room and noisy room

Error count ranks the engines that pass the blockers. In a quiet room the gaps between them are small, and noise and short clips open them up.

Shepit keeps a term list, a dictionary of the user's own names and jargon, and sends it to the engine as a hint. On the request I ship, which adds the term list and nothing more, Soniox made 0.26 errors per 100 words. That figure is a lower bound, because the rerun's scoring could not rebuild one strict layer of the reference text. A richer request, which sends extra context on top of the term list, scored 0.30 but costs more than my limit allows.

Mini's figure hides its real problem. Without the 9 clips where it had drifted before, it drops from 4.16 to 2.46, so the language drift is baked into its error count.

Soniox absolutely crushed the noisy clips! Zero errors on 15 short noisy clips, where mini made 16, or 6.58 per 100 words. Noise and short clips pulled the others apart as well:

  • ElevenLabs made 2 errors on those 15 noisy clips, GPT Transcribe 5.
  • On a 2-minute pair, quiet then noisy, Soniox made 0 and 0, ElevenLabs 2 and 11, GPT Transcribe 1 and 9.
  • On 30 short everyday clips of 1-15 s, Soniox made 3, GPT Transcribe 6, ElevenLabs 8 and mini 24.

Short clips are not an edge case, so test an engine on them.

median request from accounts other than mine
9.6 s

Seconds of speech sent per request, over 2,364 real requests from 7 accounts; my own median is 25.7 s.

English names spoken inside Ukrainian are not scored by script, because Ukrainian often writes them in Cyrillic and both spellings pass. Only Speechmatics could not write one in Latin letters at all: 0 of 28.

The accidental tap: what lands when I say nothing

An accidental tap must paste nothing. Mini never passed this check, so each survivor had to pass it on its own.

On 3 s and 8 s of room noise, mini returned its instruction prompt word for word, and with a term list loaded it returned the prompt plus the whole list. In the app, that text lands in whatever field has focus, whether that is my own term list or the middle of a chat. At least it proved the list was loaded.

GPT Transcribe returned nothing on all three silence clips. ElevenLabs returned bracketed sound tags, "noise" and "clicking", until I switched tagging off, and after that it returned nothing, 15 of 15 times.

Soniox is the odd one. Feed it 1 s of pure digital silence and it answers politely with one English word, the same on 5 of 5 repeats: "Good." with a term list, "Morning." without one. On room noise it returned nothing, 20 of 20. I accepted it, because a real microphone never produces all-zero audio: 0.0000% of the windows I measured were all-zero.

Speed: the test clock and the clock I feel

In the app, all three kept engines land text 2-3 s after I release the key, and that is the clock a user feels. The comparison rounds timed something else, one whole file per dictation. On that clock Soniox was the slowest, and its file mode never answers in less than about 1.4 s of processing.

In the app I time the wait from release to text in hand. p50 is the median, the typical wait, and p90 is the wait that only one dictation in ten exceeds.

From key release to text in hand, per engine

In the app, on shipped requests; one speaker, September 2026.

medianto the slowest tenth012345 sElevenLabs1.93 s · n 103GPT Transcribe2.47 s · n 563Soniox2.74 s · n 104
medianto the slowest tenth012345 sElevenLabs1.93 s · n 103GPT Transcribe2.47 s · n 563Soniox2.74 s · n 104

All three land the text two to three seconds after release. On the comparison's file clock, from a 5 s clip to a 6.2 min one, GPT Transcribe took 0.9 to 7.4 s, ElevenLabs 0.6 to 8.4 s and Soniox 3.2 to 11.6 s.

The seconds and the sample
EngineMedian, sSlowest tenth, sPresses
ElevenLabs1.932.85103
GPT Transcribe2.474.47563
Soniox2.743.81104

ElevenLabs is the fastest in the app. GPT Transcribe, measured on by far the largest sample, has the longest slow tail, and Soniox keeps a tighter tail than that even though its typical wait is longer.

Streaming

Several engines stream, but only Soniox streams word by word, at a flat 0.13 s after release. Without a term list it costs $0.128 per audio hour, under its rivals' batch rates, and with a list the price depends on list size and dictation length, as the next section shows.

The others stream differently. ElevenLabs realtime bills above its batch rate, and its first draft of the text arrives about 2.2 s in. GPT Transcribe realtime returns text only in whole pieces, one per pause, and loses accuracy to the provider's voice detection. If Shepit ever types while I speak, Soniox is the candidate.

The fastest engine was cut anyway: Deepgram answered in about 0.35 s, the only column it won.

Cost: when Soniox is cheaper, and when dearer

GPT Transcribe and ElevenLabs bill a flat rate per second. Soniox also bills the term list again on every request, so it wins on long dictations with a short list and loses on short dictations with a long one.

My budget was a target near $0.20 per audio hour and a hard limit of $0.30. On the requests I ship, the three kept engines land over the target, under the limit and within 11% of each other: Soniox at $0.249, ElevenLabs at $0.263 and GPT Transcribe at $0.277.

The limit did real work. It ruled out Soniox's richer request, at $0.36 per hour for files and $0.42 for streaming, and the request with the term list alone replaced it at 0.26 errors per 100 words. It also ruled out ElevenLabs realtime at about $0.36 and Deepgram Nova-2 streaming at $0.36. Speechmatics sits exactly at the limit, at a list price of $0.30.

The arithmetic

The term list is a flat charge on every request: 10 terms cost about what 2 s of GPT Transcribe does. Without a list, Soniox costs about 40% of GPT Transcribe at any length. With one, Soniox stays cheaper up to about 24 terms on a 10 s request and about 80 on a 30 s one.

[Planned block: chart line, column, light - x = seconds of speech sent per request, 0-60; y = terms in the dictionary, 0-100; one series, the break-even line (5 s: 10 terms, 10 s: 24, 15 s: 38, 20 s: 52, 25 s: 66, 30 s: 80, 35 s: 94, out of the top near 37 s); area below labelled Soniox cheaper, field above labelled GPT Transcribe cheaper; zone on x 11-45 s: the middle half of 2,364 real requests over 7 accounts; ref y = 30: the app's routing line; pins at the two medians: other accounts 9.6 s -> 21 terms, mine 25.7 s -> 65. Finding: long request with a short list, Soniox wins; short request with a long list, GPT Transcribe wins. Source: S6 s2 fit, s3 populations]

Who pays which price

Request length belongs to the user, not to me. I ramble, and other accounts dictate in short bursts, so over a whole bill Soniox costs the same as GPT Transcribe at 100 terms for me but at only 40 for them. Above 30 terms, Shepit moves Soniox to the end of the fallback order, the queue of engines it tries in turn.

ElevenLabs stays flat at $0.263 per audio hour with a term list. Every price here comes from billing records, except Deepgram Nova-3 (a free trial) and Speechmatics (list price).

Verdict: three kept, two cut

Each kept engine passed every blocker, and each wins a different column:

  • Soniox: the fewest errors, 0.26 per 100 words on the request I ship, and clean on noise.
  • ElevenLabs Scribe v2: the shortest wait in the app.
  • GPT Transcribe: a flat per-second bill and identical answers on repeats. It is the default.

I cut two engines, each on a named measure. Deepgram Nova-2 and Nova-3 lost speech on up to 6 of 76 clips and made 5.18 errors per 100 words. Speechmatics Linden 1 made 8.70 errors per 100 words, chopped sentences mid-phrase and wrote no English name in Latin letters. Both verdicts hold for Ukrainian dictation, one speaker, one microphone, one room and model versions of September 2026, and another team's audio is another test.

Routing follows from the table. GPT Transcribe is the default, and when an engine fails, Shepit re-sends the dictation to the next engine inside the same request. With three vendors, no single outage takes dictation down, and mini stays on only as the last resort.

The blockers decided who survived. The error rate only decided the order among them. Hold the key, speak, release. The words land in the field.

Written by

Oleksij Rak

Co-founder of Rebbix, maker of Shepit

Engineer for 30 years. Started Shepit to fix his own dictation and builds it hands-on, from the speech pipeline to the tools behind this blog.

Published Last verified