Talk for an hour, get your text in seconds

Quick answer

A long dictation can outgrow a speech engine's limits and leave you waiting after you stop. Shepit cuts at pauses and transcribes the pieces while you speak. In a local test, an hour assembled from real recordings produced raw text 0.9–3.2 seconds after release. Most transcription had already finished; only outstanding pieces remained. That is the wait for raw text: AI formatting of the complete transcript takes additional time.

Published
Reading time
17 min

Hold the key, speak, let go, and the text lands where your cursor is: that is how Shepit should work. Long dictations gave that sequence several ways to go wrong: an engine could refuse the recording or quietly leave part of it out, large uploads limited how many sessions my server could handle, and the wait grew with every minute spoken. Switching engines and compressing the audio each helped, but all the transcription still waited until I stopped talking.

Every engine has a ceiling

The model I was testing on, gpt-4o-mini-transcribe, had a limit on its answer: it could return only so many tokens, the small units a model uses to represent text. Audio could fit within its input limit while the transcript needed more room than its output allowed.

The request could return HTTP 200, the code for success, with part of the speech missing. I covered that hole in the middle in the engine comparison. A successful response was not proof of a complete transcript. Even checking whether the answer had reached its token cap missed some incomplete results.

gpt-transcribe moved the boundary much further out. A recording of 3,599 seconds came back complete; one of 3,600 seconds was refused with a message saying the audio might be corrupted. Re-encoding it did not help. An extra second took the recording from a complete transcript to an error.

Soniox and ElevenLabs passed the hour-long test, and their declared limits sit further out still. I had not tested those ceilings. Changing engines gave a long recording more room; sending small pieces offered a way to stay comfortably below the limits instead of finding the next one during a dictation.

Where each engine's limits sit against a recording's length

Single test files per length. Tested results drawn; Soniox and ElevenLabs declared limits lie far past the hour and were not tested.

refused0102030405060 minmini's output cap bindsgpt-4o-mini-transcribe, 1,200 saccepted, speech missing, HTTP 200gpt-4o-mini-transcribe, 1,500 srefusedgpt-transcribe, 3,599 scompletegpt-transcribe, 3,600 s1Soniox, 3,595 s2ElevenLabs, 3,595 s3
  1. 1gpt-transcribe, 3,600 s: refused: audio may be corrupted
  2. 2Soniox, 3,595 s: complete; declared limit 300 min
  3. 3ElevenLabs, 3,595 s: complete; declared limit 10 h
refused01020304060 minmini's output cap bindsgpt-4o-mini-transcribe, 1,200 s1gpt-4o-mini-transcribe, 1,500 srefusedgpt-transcribe, 3,599 s2gpt-transcribe, 3,600 s3Soniox, 3,595 s4ElevenLabs, 3,595 s5
  1. 1gpt-4o-mini-transcribe, 1,200 s: accepted, speech missing, HTTP 200
  2. 2gpt-transcribe, 3,599 s: complete
  3. 3gpt-transcribe, 3,600 s: refused: audio may be corrupted
  4. 4Soniox, 3,595 s: complete; declared limit 300 min
  5. 5ElevenLabs, 3,595 s: complete; declared limit 10 h
refused01020304060 minmini's output cap bindsgpt-4o-mini-transcribe, 1,200 s1gpt-4o-mini-transcribe, 1,500 s2gpt-transcribe, 3,599 scompletegpt-transcribe, 3,600 s3Soniox, 3,595 s4ElevenLabs, 3,595 s5
  1. 1gpt-4o-mini-transcribe, 1,200 s: accepted, speech missing, HTTP 200
  2. 2gpt-4o-mini-transcribe, 1,500 s: refused
  3. 3gpt-transcribe, 3,600 s: refused: audio may be corrupted
  4. 4Soniox, 3,595 s: complete; declared limit 300 min
  5. 5ElevenLabs, 3,595 s: complete; declared limit 10 h

In these gpt-4o-mini-transcribe tests, the output cap hit before the input ceiling, and incomplete text could still return HTTP 200. gpt-transcribe moved the boundary to one second short of an hour; the other two engines passed the hour, and their declared limits remain untested.

The limits and how each one shows
Engine, limitWhere it sitsHow it showsMeasured
gpt-4o-mini-transcribe, output cap of 2,048 tokensabout 530-670 s, by speech rateHTTP 200, part of the speech missingYes
gpt-4o-mini-transcribe, input1,200 s accepted, 1,500 s refusedHTTP 400Yes
gpt-transcribe, length3,599 s complete, 3,600 s refusedHTTP 400, audio may be corruptedYes
Soniox, declared 300 mincomplete at 3,595 s-Hour only
ElevenLabs, declared 10 hcomplete at 3,595 s-Hour only

Every minute adds to the wait

Taking the whole recording was only the first hurdle. In the whole-file approach, transcription still started after I let go. An accepted file could keep me waiting simply because there was more speech to process.

My own dictations showed a roughly linear relationship. On the same engine, for whole-file recordings longer than a minute, each extra minute spoken added about a second and a half of waiting on average. I measured that wait from releasing the key to receiving the raw text, before it was pasted. The engine could handle the recording; finishing it still took longer as the recording grew.

One six-and-a-half-minute dictation in a development build showed how bad the wait could get. The request timed out 46 seconds after release. Another request eventually returned the text, about 86 seconds after I had stopped talking. That was one failure, not the typical delay.

A separate bench test pushed whole files towards an hour. gpt-transcribe took 92 seconds on the longest file; ElevenLabs took 31. These were individual provider-call timings, not the wait through the desktop app, and each length was tested once. The faster engine made a substantial difference, but it still received the whole recording after it was finished.

Whole-file wait after release grows with length, gpt-transcribe

All 68 retained whole-file dictations over a minute; key release to raw text in hand, before paste.

dictationskey release to raw text051015 s+1.54 s per minute1234567 minrecorded length
dictationskey release to raw text051015 s+1.54 s per minute1234567 minrecorded length
dictations+1.54 s per minutekey release to raw text051015 s1234567 minrecorded length

An association across these presses, not a fixed charge per minute, fitted only over the lengths observed. One speaker, hold mode, no AI polish; the lowest and highest tenth of waits dropped per length bin; different days and network load.

The line and its sample
MeasureValue
Dictations over 60 s68
Slope, s per minute+1.54
Intercept, s2.27
R²0.88
Longest, s431

Heavy files, few sessions

The audio had to get through my server too. A minute of uncompressed WAV weighed almost two megabytes, so a long dictation became a large upload before transcription could begin. At roughly thirteen minutes, raw audio reached the speech provider's file-size limit. That was another boundary to deal with, separate from the models' limits on audio duration and output text.

I load-tested the server by replaying one real dictation, about five minutes long, from increasing numbers of simulated users. One instance processed about eight requests at once; at fifteen users, 92% of requests failed. Doubling its memory moved the throughput ceiling by exactly zero. The constraints were the amount of request data it could handle per second and the number of requests its workers could process together.

I started by compressing the recording on the Mac, before sending it. Repeating the test with the same speech encoded as AAC, a much smaller file, reached forty simulated users without a failed request and more than quadrupled throughput. I stopped on the test's budget, not on the server: the new ceiling had not been found. Compression was a useful first fix, but the engine still received the whole recording only after I let go.

Compressing the same speech took one server instance from failing at 15 users to clean at 40.

One real 289.6 s dictation replayed by simulated users against one instance, 1 vCPU, 512 MiB.

MeasureRaw WAVCompressed AAC
File size9.27 MB907 KB, 24 kbps
Capacity ceiling, one instanceabout 8 served (corrected by later calibration)ceiling not reached at 40 offered
At 15 users92.4% failed (160 of 2,097 succeeded)-
At 40 users-0 failed of 240
Whole 10-40 user ladder-0 failed of 1,151
Throughput, dictations / min55254
Doubling memory to 1 GiBno throughput gain-

The compressed run stopped at the test's budget, not at a failure, so 40 users is not its ceiling. The app now ships 32 kbps, about 8 times smaller than raw; no load test exists after cutting.

Start before I finish

I was sending one whole recording after I stopped talking. Compression made it lighter, but none of its transcription happened while I spoke. Cutting offered a way to use that time: send a finished piece before the rest of the recording exists.

Each piece goes to the engine in an ordinary file request while I record the next one. The client keeps the returned text in recording order. Most transcription can finish before I let go; after release, the app waits for whatever is still outstanding, usually the final piece, then joins the text. The remaining work depends on those outstanding pieces rather than a whole recording that has only just been sent.

Transcribe each finished piece while I am still recording the next one.

The cutter's two-minute ceiling keeps each piece comfortably below the engines' length limits and keeps its audio upload small. Smaller bodies should also ease the byte-throughput constraint from the load tests, although I have no post-cutting capacity measurement. This addressed the request shape behind several problems at once. The next problem was the join: short files would not help much if splitting a word or a sentence made their transcripts harder to put back together.

First try: overlap and stitch

My earlier experiments used fixed-length pieces, so a cut could land in the middle of a word. I repeated a little audio on both sides of the cut to give the engine another chance at that speech. This returned two versions of the same passage. Joining them meant finding where the texts agreed and keeping the passage once.

The unfinished matcher searched too far back in the earlier text. Two common words in the wrong place convinced it that it had found the join, and it silently deleted 79 words of real speech, with no error raised. This time the missing words came from my own joining code. Restricting the search to the end of one piece and the beginning of the next repaired it. Across four long Ukrainian dictations, the bounded matcher showed a net word-count difference of four on 1,505 words, with no invented word sequences detected and every join reading coherently. It was a useful result from that tested set.

I also tried having a language model do the joining. Given the whole text, it produced visible duplication at the joins. Restricting it to the joins worked much better, but still deleted three real words and added paid model calls. The bounded algorithm did the job better. I had a workable way to join overlapping text, but every cut still created something to reconcile.

Four ways to join overlapping pieces, and what each did to the text

Earlier experiments on gpt-4o-mini-transcribe with overlapping pieces; each row has its own measure, so the rows share no common score.

MethodObserved resultConditionsExtra cost per audio hourOutcome
Bounded matchernet word count -4 of 1,505; no invented word sequences; 8 of 8 joins coherent4 longest Ukrainian dictations, 90 s pieces, 10 s overlap+3.86% uk, +3.05% enWorked, never shipped
Unbounded matcher79 real words deleted, silentlya false two-word match 46 words back; unfinished prototype-Fixed by bounding
Language model, whole textvisible duplication in 4 of 6 joinsgpt-4.1-mini-Dropped
Language model, joins only3 real words deletedgpt-4.1-mini on the join regions+10.71% uk, +11.19% enWorse than bounded

A net word-count difference is not four established omissions, nor an error score against a human reference. Costs compare with one whole-file call. The shipped cutter makes no overlap: none of these matchers ships.

Silence changes the question

The overlap was there because a fixed-length cut could land in the middle of a word. At a pause between phrases, there is no word to split: the pieces join end to end, in order, with no audio sent twice and nothing to match. My silence measurements showed such pauses are common: almost half the pauses between phrases lasted more than a second.

To hear a pause while I am still talking, the cutter uses the same kind of live detector as the recording indicator, with its own settings. It should call silence only when it is sure: a missed pause just makes a piece a little longer, while a false pause in the middle of a word cuts the word in two.

Cutting has a price of its own: each piece is transcribed without hearing the rest of the recording. I compared the text after cutting and trimming the pauses with the text of the same recordings sent whole, counting changed, added or missing words. There were 1.16 such differences per 100 words on average. Repeating the whole-file request produced 0.28 differences per 100 words. I did not check either transcript against what I actually said, so this comparison cannot tell which version got a word right. Cutting at pauses made the join simple; it did not give the engine back the rest of the recording.

How long should a piece be

My first rule waited forty seconds, then cut at the next pause. It had no ceiling, so nothing in the rule bounded a piece when a suitable pause failed to arrive. Later replays showed a forty-second start made many more joins with no measured gain, so the search moved to start after a minute, with a two-minute ceiling. These were search limits, not a schedule for cutting at exact intervals. Most short dictations stayed untouched: 88% of the tested archive went to the engine as one piece.

Then I looked at where the cuts actually landed. The typical piece lasted seventy-two seconds, not sixty, so I asked whether the floor could move down until the median cut fell on the minute. The study replayed 1,046 dictations and answered with a price. A forty-eight-second floor produced 44% more joins and 63% more final pieces shorter than fifteen seconds, for about 0.2 seconds less predicted median wait after release. That wait was modelled from the final pieces of recordings over two minutes, not newly measured through the app. That was a bad trade, and the one-minute start stayed.

Moving the floor later was not a clear win either. It reduced joins but increased the modelled worst wait as longer pieces became possible. The study recommended keeping both, with medium confidence. I had wanted smaller pieces, but it persuaded me: in my own use that balance is working well, and I am not planning to change it.

Moving the floor earlier buys many more joins for a fraction of a second; 60 / 120 stayed.

Replay of 1,046 dictations (11.6 h, one speaker). Joins and short final pieces: whole archive. Waits: modelled on the 55 recordings over 120 s.

Floor / ceiling, sJoinsFinal pieces under 15 sModelled wait, median sModelled wait, worst s
40 / 90289651.592.68
48 / 108228521.572.97
60 / 120, chosen158321.782.97
75 / 135108221.783.80
90 / 15083171.734.10

Waits come from a model of the final piece (604 ms plus 23.7 ms per second of it), not from the app. The replay ran the live cutter before its voice look-back was added, and made no new engine calls.

Hearing a pause while you talk

The cutter has only the audio already recorded. Once a piece passes a minute, it waits for the first pause that begins after that point and lasts a full second. When the second is confirmed, it puts the boundary about halfway into the pause. The decision comes after the boundary, without needing to hear what I will say next.

At two minutes, it stops waiting and looks back over that piece's search window. First it chooses the longest pause of at least half a second. If the gate found none, the trimmer's voice detector gets one look at the recorded audio, looking for a non-speech stretch of at least a fifth of a second in the same window. I added that step after the archive replay produced two hard cuts in 158 cuts. In one recording, the live gate missed an earlier short gap; the voice scan moved the boundary from 119.9 seconds back to 82.1, out of the stretch of speech.

If even that scan finds no suitable boundary, the cutter uses the quietest point in the last ten seconds. That hard cut keeps the piece bounded, but can still divide speech. The gate can miss a real pause: in the unusually long waits I examined, sound above its silence line or a short burst inside a pause often stopped it accepting silence. Waiting longer was not always evidence that I had spoken without stopping.

Where the cutter cut one real two-minute dictation

Each piece's search window opens 60 s after its own start. The cutter cuts about 0.51 s into a pause, once it has confirmed the pause lasts a full second.

where the next cut may landa piece, sent while you talkthe last piece, sent at releasewindow left unusedcut at 62.27 spiece 1 ceiling, if no pause is found1the recordingthe possible search windowsunused: piece 1 already cutthe piecespiece 1piece 20306090120 s
  1. 1cut at 124.99 s
where the next cut may landa piece, sent while you talkthe last piece, sent at releasewindow left unusedcut at 62.27 s1the recording2the possible search windowsunused: piece 1 already cutthe piecespiece 1piece 20306090120 s
  1. 1piece 1 ceiling, if no pause is found
  2. 2cut at 124.99 s
where the next cut may landa piece, sent while you talkthe last piece, sent at releasewindow left unusedcut at 62.27 s1the recording2the possible search windows3the piecespiece 1piece 20306090120 s
  1. 1piece 1 ceiling, if no pause is found
  2. 2cut at 124.99 s
  3. 3unused: piece 1 already cut

Both cuts found a full-second pause soon after their window opened, so neither piece reached its two-minute ceiling. Piece 2's window is counted from the first cut, not from a fixed two-minute grid.

The cuts, in seconds
PieceWindow opensCut
16062.27
2122.27124.99
last-release, 126.4

Fast, whatever the length

The change shows up where I wanted it: after I let go of the key. The chart plots my hold-to-talk recordings against the wait for raw text. Below a minute, both clouds still climb because nothing has been cut. Beyond a minute, whole-file recordings added about 1.7 seconds of wait per extra minute spoken on average across engines, and about 1.5 seconds on gpt-transcribe alone. Recordings sent in pieces stayed roughly flat over the lengths observed.

Wait after release against recorded length, before and after cutting

Showing 426 of 3,077 retained dictations: up to 16 actual recordings per engine, era and length bin. Dot density is not recording frequency.

whole filein pieceskey release to raw text105101520 swhole file: +1.67 s/minin pieces: flat012345678910 minrecorded length
  1. 1cutting starts after a minute
whole file: +1.67 s/minin pieces: flatwhole filein pieceskey release to raw text105101520 s012345678910 minrecorded length
  1. 1cutting starts after a minute
whole file: +1.67 s/minin pieces: flatwhole filein pieceskey release to raw text105101520 s01234567810 minrecorded length
  1. 1cutting starts after a minute

Each cloud mixes engines (whole file mostly gpt-4o-mini, pieces mostly Soniox and ElevenLabs); per-engine slopes below. Fits: all retained recordings over 60 s per era, over observed lengths. Retained: lowest and highest tenth of waits dropped per era, engine, length bin; AI polish, paste excluded. One speaker, room and language; hold mode; different weeks; network and provider load uncontrolled.

The counts and the lines behind the picture
Era, engineRetainedShownOver 60 sOver 60 s, s/minLongest, s
Whole file, all engines1,902175371+1.67431
In pieces, all engines1,175251280-0.11608
Whole file, gpt-4o-mini-transcribe1,59791303+1.71373
Whole file, gpt-transcribe3058468+1.54431
In pieces, gpt-transcribe1717635-0.01248
In pieces, Soniox68090170-0.12608
In pieces, ElevenLabs3248575-0.03425

The two eras have different engine mixes: most earlier recordings used gpt-4o-mini-transcribe, while most later ones used Soniox or ElevenLabs. The shared gpt-transcribe comparison shows the same turn, and each cut engine has a near-flat line of its own above a minute. Before cutting, my five-to-ten-minute whole-file dictations across engines waited a median of about ten seconds, the worst nearly half a minute.

The cuts also have a live check in my production and development logs: 323 of 331 landed at a full pause, the remainder at shorter ones, and none required a hard cut. My longest ordinary dictation so far was a thought I dictated for a little over ten minutes. Soniox handled it in eight pieces, and the raw text was in hand 4.4 seconds after release. Dictations that long are still rare for me.

raw text after release, a ten-minute dictation in eight pieces
4.4 s
cuts that landed at a full-second pause
323 of 331
hard cuts
0 of 331

4.4 s: one Soniox dictation of a little over ten minutes, before paste. Cuts: 234 cut dictations in my production and development logs; observations, not a guarantee.

An hour on the bench

My real dictations had reached about ten minutes, so testing an hour meant assembling one. I joined 124 real recordings into just under an hour of audio, with room noise between them, and ran it through the desktop's recording code with gpt-transcribe and a local API. In two raw-text runs, it made 44 pieces. At speaking speed, only a short final piece remained after release; when I fed the audio ten times faster, an earlier piece was still in flight too. Raw text was ready in 0.9-3.2 seconds across the runs, depending on the outstanding pieces. A similarly long file took 92 seconds in the separate whole-file provider test. Most of the piecewise pipeline's work had happened while the recording was still going.

That timing does not certify a complete transcript. The test used an earlier trimmer, and 122 of the 124 clips came back in the text. It has not been repeated on the current build. Both missing clips were the same quiet whisper; one of them the engine loses even untrimmed. The silence-trimming article covers the trimmer's part. These raw-text timings exclude the production network hop and paste. AI formatting of the whole text adds its own wait, covered below.

What the pieces open up

While I am still speaking, the app already holds text from the pieces that have returned, in order, on my Mac. Today the pieces stay invisible; the user still gets one dictation. Having that text early makes room for a future interface to show progress through a long recording, based on work already finished.

These are ordinary file requests, rather than a realtime transcription session. In my comparison, realtime cost more on Soniox, ElevenLabs and Deepgram, though not on gpt-transcribe; Soniox's realtime mode also bills the whole session, silence included, so trimming pauses would save nothing there. Cutting let me keep the file engines and finish most of their work before release. In my longer dictations, I can keep following the thought, then let go and have the raw text in a few seconds.

Written by

Oleksij Rak

Co-founder of Rebbix, maker of Shepit

Engineer for 30 years. Started Shepit to fix his own dictation and builds it hands-on, from the speech pipeline to the tools behind this blog.

Published