How I chose a speech-to-text engine
Six engines, 1,000+ of my own dictations, one noisy room. A dictation in the wrong language or speech that never arrived disqualified an engine before its error count was read.
Quick answer
The engine with the fewest errors did not become my default. I tested six speech-to-text engines on my own dictations in a noisy room, and the worst failures were not wrong words. Some engines dropped parts of what I said, and my old model sometimes returned a whole dictation in another language. I kept three, Soniox, ElevenLabs Scribe v2 and OpenAI GPT Transcribe, and made GPT Transcribe, with its flat per-second bill and identical answers on repeats, the default.
I tested six speech-to-text engines on 1,000+ of my own Ukrainian dictations, about 10 hours of speech. The one with the fewest errors did not become my default.
The model Shepit had been running on faced Soniox, ElevenLabs Scribe v2 and OpenAI GPT Transcribe. Deepgram Nova-2 and Nova-3 and Speechmatics Linden 1 joined them. Three stayed, two were cut, and the old model dropped to last resort. The error rate ranked the survivors, but it did not choose them.
If you are choosing a speech-to-text API for a language other than English, this test is for you, because the error rate hides the failures that matter.
Why I looked
Shepit has one job: hold the key, speak, release, and the words land in the field. The engine under that job was OpenAI gpt-4o-mini-transcribe, mini from here on, and mostly it did the job. Then I read my own transcripts properly, and four problems kept turning up.
It changed language on me. I dictate in Ukrainian, yet on 8 of the 76 test clips the whole dictation came back in another language, not just a word. Pinning the language in the request did not help: on the clips where mini had drifted before, 37 of 45 repeats drifted again.
It transcribed silence. Tap the hotkey by accident, say nothing, and mini turns 3 s of room noise into text: Shepit's own instruction prompt.
It never answered the same way twice. I ran 9 clips 5 times each, and not one of the 9 came back identical across its repeats. Soniox, GPT Transcribe and Deepgram gave identical text on 9 of 9.
It cut the middle out of long recordings. A 900 s file came back missing everything from 200 s to 660 s, so only 36% of the words arrived. HTTP 200, no error, no apology. The first and last lines looked fine, so a glance at the ends would never catch it.
The old model failed in ways a user cannot see or fix, not in how many words it misheard.
None of these is a wrong word. You see a wrong word and fix it, while a dictation in the wrong language, a transcript of nothing or a hole in the middle turn up later, if at all.
So I did not go into the test asking which engine mishears least. I asked which engine never does any of this, and only then, among the survivors, which one mishears least.
The scoreboard: blockers first, error count second
I wrote the rules down before I saw a single score, so I could not argue with them later. Two failures disqualify an engine outright: a whole dictation returned in the wrong language, and speech that never arrives. Both are worse than a wrong word, which at least shows up on screen. One more check applies: the engine must not invent text when I say nothing.
The rest of my brief was short:
- No language drift and no dropped speech: critical.
- Streaming, text that appears while I speak: very important.
- Price: a target near $0.20 per audio hour, a hard limit of $0.30.
A blocker disqualified an engine, whatever its error rate.
76 of my hardest Ukrainian dictations; errors per 100 words after my rulings. Coloured: the cell that cut or replaced the engine.
| Engine | Dictations in another language, of 76 | Clips with lost speech, of 76 | Silence check | Errors / 100 words | Verdict |
|---|---|---|---|---|---|
| Soniox | 0 | 0 | Yes, one word on digital silence | 0.26 | Kept |
| ElevenLabs Scribe v2 | 0 | 0 | Yes, with sound tags off | 1.29 | Kept |
| OpenAI GPT Transcribe | 0 | 1 | Yes, empty | 1.59 | Kept, the default |
| OpenAI gpt-4o-mini-transcribe | 8 | 1 | No, returned its own prompt | 4.16 | Replaced, last resort |
| Deepgram | 0 | 6 | Yes, empty | 5.18 | Cut: lost speech |
| Speechmatics | 0 | 0 | Yes, empty | 8.70 | Cut: error count, chopped sentences |
Soniox shows the file-mode request I ship; every other row is the engine's best setup: ElevenLabs Scribe v2 with Ukrainian set and sound tags removed, GPT Transcribe in file mode, mini as Shepit shipped it, Deepgram Nova-3 on finished files and Speechmatics Linden 1 streaming, all but mini with my list of 54 names and terms, of which Soniox's shipped request sends the first 50. Soniox's 0.26 is a lower bound, and its best setup, with extra context, scored 0.30 but costs $0.36 per audio hour, over my limit.
Read the table by column, not by row. With the language locked, Deepgram and Speechmatics never switched language once, and I still cut both.
Soniox made the fewest errors, and that alone did not buy it a pass: it still had to clear the silence check before I kept it. GPT Transcribe, the engine I made the default, sits only third in the error column.
How I tested
Every clip in this test is mine: real dictations from real use, not a synthetic benchmark. AI agents made the scale possible, and I kept the last word.
- dictations in the corpus
- 1,000+
- of my speech in the corpus
- 10 h
- transcriptions
- 4,442+
- of audio transcribed
- 38.6 h
- of agents
- dozens
One speaker, one microphone, one room; dictations of 14-22 September 2026; September 2026 model versions.
The 4,442+ is a lower bound, because refused and retried calls are not counted. It is the sum of four parts:
- 3,543 in the two comparison rounds, 1,200 of them streaming sessions
- 635 on the requests I ship
- 243 in a check of my list of names and terms
- 21 files of growing length, up to an hour each
The clips
The dictations in the corpus were recorded from 14 to 22 September 2026. I screened 1,000 of them and picked 71: the clips where mini had failed or was most likely to, and a control group.
- 9 it had already switched language on
- the 20 noisiest
- 20 short ones
- 12 with English names inside Ukrainian
- 10 ordinary clips as the control
I added 5 long recordings, up to 6.2 minutes, for 76 test clips in all, and put 3 silence clips beside them. "Noisy" is measured, not felt: those 20 clips have the lowest signal-to-noise ratio, the gap between my voice and the room's hum.
The 76 are where all six engines met head to head. Beyond them I also checked the requests I ship, files up to an hour long and the app's own timings.
What counts as an error
The usual yardstick is word error rate (WER). Mine counts only content words, the ones that carry meaning, heard wrong, dropped or invented, and ignores small grammar words, word endings and spelling. Every rate below is errors per 100 words of the reference text, the correct transcript each engine is scored against.
The reference text
I built my first reference text from mini's own output and checked every disputed word against Soniox. On that reference, Soniox's best setup scored 0.04 errors per 100 words. Suspiciously close to perfect: my own experiment had slipped it part of the answer key. So I ruled 17 disputed clips myself and set 4 global rules, and Soniox's best setup went from 0.04 to 0.30 while mini's went from 3.22 to 4.16.
The agents and I
AI agents built and ran the pipeline, from picking the clips to scoring them. Dozens of them worked with me for about a day and a night: they researched by day, ran experiments overnight and brought verdicts in the morning. They scored 3,502 outputs from the two comparison rounds the same way and surfaced 127 disputed words.
The last word stayed mine. My 17 clip rulings and 4 global rules override the agents everywhere, and Speechmatics' 8.70 outlier got a spot check by hand.
Blocker: the wrong language
A whole dictation in another language is a broken dictation, not one wrong word. No keystroke fixes it: you delete it and speak again, if you notice.
Mini's switches were no accident of the test. 7 of its 8 landed on the 9 clips where it had already drifted in real use, so the same clips failed the same way.
- whole dictations in another language
8then 0 of 76
The old model against every kept engine, best setups, the language hint sent.
The two cut engines switched no dictation wholesale either, and the kept three hold that on the requests I ship too. What remains is one Russian-looking word form on 1 or 2 clips, scored as a single error.
The language hint carries real weight. Without its language code, ElevenLabs switched 6 of 44 clips wholesale, Croatian and Slovak among them, though I have never dictated a word of either. With the code sent, 76 of 76 came back Ukrainian. Mini is the exception: the same hint did not stop its drift.
GPT Transcribe has two modes. File mode transcribes the finished recording, and realtime mode transcribes while I speak, cutting the audio into pieces wherever the provider's voice detection hears a pause. That cutting brings the drift back: in realtime mode it switched 1 clip wholesale, while in file mode the same model switched none.
Blocker: speech that never arrives
Lost speech is the worst failure, because nobody notices it. A wrong word is on screen. A missing sentence is not.
Deepgram's best setup lost speech on 6 of 76 clips: 1 came back empty, 2 missed the start and 3 missed the end. Plain Nova-2 returned 3 clips empty. Its language never drifted, and I still cut it, on Ukrainian dictation, one speaker, one room and September 2026 model versions.
Realtime voice detection has its own way to lose a dictation. On a very quiet recording, at -71 dBFS, far below normal speech level, the provider's detection never fired. The audio was never sent on for transcription, and no error came back. The whole recording was gone.
The worst form is a hole in the middle, and mini left one in two long files. Both times it answered HTTP 200 with no signal. A hell of a success code. Shepit's own truncation check, which watches the length of mini's answer, misses such cases. If you test a speech-to-text API, check the middle of long files.
What came back from long recordings: a hole in the middle for mini, the whole file for the kept three
My files of growing length, up to 3,595 s; September 2026 model versions. Coloured: the speech that never arrived. Bold: the whole file came back.
Mini's text begins where the recording begins and ends where it ends, so a reader who checks the first and last lines sees nothing wrong: the hole is in the middle. At 1,500 s it refused the file outright, while the three kept engines returned 3,595 s complete.
The files and what came back
| Engine, file | Text came back, s | Speech missing, s | Weakest minute, share of words | Outcome |
|---|---|---|---|---|
| mini, 900 s | 0-200, 660-900 | 200-660 | - | Hole in the middle |
| mini, 1,200 s | 0-420, 1,140-1,200 | 420-1,140 | - | Hole in the middle |
| mini, 1,500 s | none | - | - | Refused |
| Soniox, 3,595 s | 0-3,595 | - | 0.96 | Complete |
| ElevenLabs, 3,595 s | 0-3,595 | - | 0.95 | Complete |
| GPT Transcribe, 3,595 s | 0-3,595 | - | 0.84 | Complete |
The kept three passed. Soniox and ElevenLabs lost speech on 0 of 76 clips, and GPT Transcribe on 1, a cut tail.
Then I fed them files of growing length, up to an hour minus five seconds, and all three came back complete. Their weakest minute, the one where the fewest words came back, fell at a different point in the file for each engine.
Error count: quiet room and noisy room
Error count ranks the engines that pass the blockers. In a quiet room the gaps between them are small, and noise and short clips open them up.
Shepit keeps a term list, a dictionary of the user's own names and jargon, and sends it to the engine as a hint. On the request I ship, which adds the term list and nothing more, Soniox made 0.26 errors per 100 words. That figure is a lower bound, because the rerun's scoring could not rebuild one strict layer of the reference text. A richer request, which sends extra context on top of the term list, scored 0.30 but costs more than my limit allows.
Mini's figure hides its real problem. Without the 9 clips where it had drifted before, it drops from 4.16 to 2.46, so the language drift is baked into its error count.
Soniox absolutely crushed the noisy clips! Zero errors on 15 short noisy clips, where mini made 16, or 6.58 per 100 words. Noise and short clips pulled the others apart as well:
- ElevenLabs made 2 errors on those 15 noisy clips, GPT Transcribe 5.
- On a 2-minute pair, quiet then noisy, Soniox made 0 and 0, ElevenLabs 2 and 11, GPT Transcribe 1 and 9.
- On 30 short everyday clips of 1-15 s, Soniox made 3, GPT Transcribe 6, ElevenLabs 8 and mini 24.
Short clips are not an edge case, so test an engine on them.
- median request from accounts other than mine
- 9.6 s
Seconds of speech sent per request, over 2,364 real requests from 7 accounts; my own median is 25.7 s.
English names spoken inside Ukrainian are not scored by script, because Ukrainian often writes them in Cyrillic and both spellings pass. Only Speechmatics could not write one in Latin letters at all: 0 of 28.
The accidental tap: what lands when I say nothing
An accidental tap must paste nothing. Mini never passed this check, so each survivor had to pass it on its own.
On 3 s and 8 s of room noise, mini returned its instruction prompt word for word, and with a term list loaded it returned the prompt plus the whole list. In the app, that text lands in whatever field has focus, whether that is my own term list or the middle of a chat. At least it proved the list was loaded.
GPT Transcribe returned nothing on all three silence clips. ElevenLabs returned bracketed sound tags, "noise" and "clicking", until I switched tagging off, and after that it returned nothing, 15 of 15 times.
Soniox is the odd one. Feed it 1 s of pure digital silence and it answers politely with one English word, the same on 5 of 5 repeats: "Good." with a term list, "Morning." without one. On room noise it returned nothing, 20 of 20. I accepted it, because a real microphone never produces all-zero audio: 0.0000% of the windows I measured were all-zero.
Speed: the test clock and the clock I feel
In the app, all three kept engines land text 2-3 s after I release the key, and that is the clock a user feels. The comparison rounds timed something else, one whole file per dictation. On that clock Soniox was the slowest, and its file mode never answers in less than about 1.4 s of processing.
In the app I time the wait from release to text in hand. p50 is the median, the typical wait, and p90 is the wait that only one dictation in ten exceeds.
From key release to text in hand, per engine
In the app, on shipped requests; one speaker, September 2026.
All three land the text two to three seconds after release. On the comparison's file clock, from a 5 s clip to a 6.2 min one, GPT Transcribe took 0.9 to 7.4 s, ElevenLabs 0.6 to 8.4 s and Soniox 3.2 to 11.6 s.
The seconds and the sample
| Engine | Median, s | Slowest tenth, s | Presses |
|---|---|---|---|
| ElevenLabs | 1.93 | 2.85 | 103 |
| GPT Transcribe | 2.47 | 4.47 | 563 |
| Soniox | 2.74 | 3.81 | 104 |
ElevenLabs is the fastest in the app. GPT Transcribe, measured on by far the largest sample, has the longest slow tail, and Soniox keeps a tighter tail than that even though its typical wait is longer.
Streaming
Several engines stream, but only Soniox streams word by word, at a flat 0.13 s after release. Without a term list it costs $0.128 per audio hour, under its rivals' batch rates, and with a list the price depends on list size and dictation length, as the next section shows.
The others stream differently. ElevenLabs realtime bills above its batch rate, and its first draft of the text arrives about 2.2 s in. GPT Transcribe realtime returns text only in whole pieces, one per pause, and loses accuracy to the provider's voice detection. If Shepit ever types while I speak, Soniox is the candidate.
The fastest engine was cut anyway: Deepgram answered in about 0.35 s, the only column it won.
Cost: when Soniox is cheaper, and when dearer
GPT Transcribe and ElevenLabs bill a flat rate per second. Soniox also bills the term list again on every request, so it wins on long dictations with a short list and loses on short dictations with a long one.
My budget was a target near $0.20 per audio hour and a hard limit of $0.30. On the requests I ship, the three kept engines land over the target, under the limit and within 11% of each other: Soniox at $0.249, ElevenLabs at $0.263 and GPT Transcribe at $0.277.
The limit did real work. It ruled out Soniox's richer request, at $0.36 per hour for files and $0.42 for streaming, and the request with the term list alone replaced it at 0.26 errors per 100 words. It also ruled out ElevenLabs realtime at about $0.36 and Deepgram Nova-2 streaming at $0.36. Speechmatics sits exactly at the limit, at a list price of $0.30.
The arithmetic
The term list is a flat charge on every request: 10 terms cost about what 2 s of GPT Transcribe does. Without a list, Soniox costs about 40% of GPT Transcribe at any length. With one, Soniox stays cheaper up to about 24 terms on a 10 s request and about 80 on a 30 s one.
[Planned block: chart line, column, light - x = seconds of speech sent per request, 0-60; y = terms in the dictionary, 0-100; one series, the break-even line (5 s: 10 terms, 10 s: 24, 15 s: 38, 20 s: 52, 25 s: 66, 30 s: 80, 35 s: 94, out of the top near 37 s); area below labelled Soniox cheaper, field above labelled GPT Transcribe cheaper; zone on x 11-45 s: the middle half of 2,364 real requests over 7 accounts; ref y = 30: the app's routing line; pins at the two medians: other accounts 9.6 s -> 21 terms, mine 25.7 s -> 65. Finding: long request with a short list, Soniox wins; short request with a long list, GPT Transcribe wins. Source: S6 s2 fit, s3 populations]
Who pays which price
Request length belongs to the user, not to me. I ramble, and other accounts dictate in short bursts, so over a whole bill Soniox costs the same as GPT Transcribe at 100 terms for me but at only 40 for them. Above 30 terms, Shepit moves Soniox to the end of the fallback order, the queue of engines it tries in turn.
ElevenLabs stays flat at $0.263 per audio hour with a term list. Every price here comes from billing records, except Deepgram Nova-3 (a free trial) and Speechmatics (list price).
Verdict: three kept, two cut
Each kept engine passed every blocker, and each wins a different column:
- Soniox: the fewest errors, 0.26 per 100 words on the request I ship, and clean on noise.
- ElevenLabs Scribe v2: the shortest wait in the app.
- GPT Transcribe: a flat per-second bill and identical answers on repeats. It is the default.
I cut two engines, each on a named measure. Deepgram Nova-2 and Nova-3 lost speech on up to 6 of 76 clips and made 5.18 errors per 100 words. Speechmatics Linden 1 made 8.70 errors per 100 words, chopped sentences mid-phrase and wrote no English name in Latin letters. Both verdicts hold for Ukrainian dictation, one speaker, one microphone, one room and model versions of September 2026, and another team's audio is another test.
Routing follows from the table. GPT Transcribe is the default, and when an engine fails, Shepit re-sends the dictation to the next engine inside the same request. With three vendors, no single outage takes dictation down, and mini stays on only as the last resort.
The blockers decided who survived. The error rate only decided the order among them. Hold the key, speak, release. The words land in the field.
Questions I get asked
Which speech-to-text API is most accurate for Ukrainian?
Soniox, in my test: 0.26 errors per 100 words, against 1.29 for ElevenLabs Scribe v2 and 1.59 for OpenAI GPT Transcribe. That holds for 76 of my hardest Ukrainian dictations, one speaker, one room and September 2026 model versions, with my list of names and terms sent as a hint. Soniox also made zero errors on 15 short noisy clips. I still made GPT Transcribe the default, for its flat per-second bill and identical answers on repeats.
Why does speech-to-text switch to another language, and how do I stop it?
Send the language code with every request. Without it, ElevenLabs Scribe v2 switched 6 of 44 of my Ukrainian clips wholesale, Croatian and Slovak among them, and with it 76 of 76 came back Ukrainian. The code is not a cure for every model: OpenAI gpt-4o-mini-transcribe kept drifting with it, on 37 of 45 repeats of the clips where it had drifted before. Realtime mode brings the risk back, because voice detection cuts the audio at every pause: OpenAI GPT Transcribe switched 1 clip in realtime mode and none in file mode.
Is Soniox cheaper than OpenAI GPT Transcribe?
It depends on the term list and the request length, because Soniox bills the term list again on every request. Without a list, Soniox costs about 40% of GPT Transcribe at any length. With one, it stays cheaper up to about 24 terms on a 10 s request and about 80 on a 30 s one. In my September 2026 test, on the requests Shepit sends, Soniox came to $0.249 per audio hour, ElevenLabs Scribe v2 to $0.263 and GPT Transcribe to $0.277.
How do I test a speech-to-text API before choosing one?
Run it on your own recordings and check what the error rate hides. Check the middle of long files: OpenAI gpt-4o-mini-transcribe answered HTTP 200 on a 900 s file with everything from 200 s to 660 s missing, while the first and last lines looked fine. Test short clips, because they pull engines apart: on 30 everyday clips of 1-15 s, Soniox made 3 errors and gpt-4o-mini-transcribe 24. Feed it silence too, since an accidental tap must paste nothing, and gpt-4o-mini-transcribe turned 3 s of room noise into its own instruction prompt.