Talk for an hour, get your text in seconds
Quick answer
A long dictation can outgrow a speech engine's limits and leave you waiting after you stop. Shepit cuts at pauses and transcribes the pieces while you speak. In a local test, an hour assembled from real recordings produced raw text 0.9–3.2 seconds after release. Most transcription had already finished; only outstanding pieces remained. That is the wait for raw text: AI formatting of the complete transcript takes additional time.
Hold the key, speak, let go, and the text lands where your cursor is: that is how Shepit should work. Long dictations gave that sequence several ways to go wrong: an engine could refuse the recording or quietly leave part of it out, large uploads limited how many sessions my server could handle, and the wait grew with every minute spoken. Switching engines and compressing the audio each helped, but all the transcription still waited until I stopped talking.
Every engine has a ceiling
The model I was testing on, gpt-4o-mini-transcribe, had a limit on its answer: it could return only so many tokens, the small units a model uses to represent text. Audio could fit within its input limit while the transcript needed more room than its output allowed.
The request could return HTTP 200, the code for success, with part of the speech missing. I covered that hole in the middle in the engine comparison. A successful response was not proof of a complete transcript. Even checking whether the answer had reached its token cap missed some incomplete results.
gpt-transcribe moved the boundary much further out. A recording of 3,599 seconds came back complete; one of 3,600 seconds was refused with a message saying the audio might be corrupted. Re-encoding it did not help. An extra second took the recording from a complete transcript to an error.
Soniox and ElevenLabs passed the hour-long test, and their declared limits sit further out still. I had not tested those ceilings. Changing engines gave a long recording more room; sending small pieces offered a way to stay comfortably below the limits instead of finding the next one during a dictation.
Where each engine's limits sit against a recording's length
Single test files per length. Tested results drawn; Soniox and ElevenLabs declared limits lie far past the hour and were not tested.
- 1gpt-transcribe, 3,600 s: refused: audio may be corrupted
- 2Soniox, 3,595 s: complete; declared limit 300 min
- 3ElevenLabs, 3,595 s: complete; declared limit 10 h
- 1gpt-4o-mini-transcribe, 1,200 s: accepted, speech missing, HTTP 200
- 2gpt-transcribe, 3,599 s: complete
- 3gpt-transcribe, 3,600 s: refused: audio may be corrupted
- 4Soniox, 3,595 s: complete; declared limit 300 min
- 5ElevenLabs, 3,595 s: complete; declared limit 10 h
- 1gpt-4o-mini-transcribe, 1,200 s: accepted, speech missing, HTTP 200
- 2gpt-4o-mini-transcribe, 1,500 s: refused
- 3gpt-transcribe, 3,600 s: refused: audio may be corrupted
- 4Soniox, 3,595 s: complete; declared limit 300 min
- 5ElevenLabs, 3,595 s: complete; declared limit 10 h
In these gpt-4o-mini-transcribe tests, the output cap hit before the input ceiling, and incomplete text could still return HTTP 200. gpt-transcribe moved the boundary to one second short of an hour; the other two engines passed the hour, and their declared limits remain untested.
The limits and how each one shows
| Engine, limit | Where it sits | How it shows | Measured |
|---|---|---|---|
| gpt-4o-mini-transcribe, output cap of 2,048 tokens | about 530-670 s, by speech rate | HTTP 200, part of the speech missing | Yes |
| gpt-4o-mini-transcribe, input | 1,200 s accepted, 1,500 s refused | HTTP 400 | Yes |
| gpt-transcribe, length | 3,599 s complete, 3,600 s refused | HTTP 400, audio may be corrupted | Yes |
| Soniox, declared 300 min | complete at 3,595 s | - | Hour only |
| ElevenLabs, declared 10 h | complete at 3,595 s | - | Hour only |
Every minute adds to the wait
Taking the whole recording was only the first hurdle. In the whole-file approach, transcription still started after I let go. An accepted file could keep me waiting simply because there was more speech to process.
My own dictations showed a roughly linear relationship. On the same engine, for whole-file recordings longer than a minute, each extra minute spoken added about a second and a half of waiting on average. I measured that wait from releasing the key to receiving the raw text, before it was pasted. The engine could handle the recording; finishing it still took longer as the recording grew.
One six-and-a-half-minute dictation in a development build showed how bad the wait could get. The request timed out 46 seconds after release. Another request eventually returned the text, about 86 seconds after I had stopped talking. That was one failure, not the typical delay.
A separate bench test pushed whole files towards an hour. gpt-transcribe took 92 seconds on the longest file; ElevenLabs took 31. These were individual provider-call timings, not the wait through the desktop app, and each length was tested once. The faster engine made a substantial difference, but it still received the whole recording after it was finished.
Whole-file wait after release grows with length, gpt-transcribe
All 68 retained whole-file dictations over a minute; key release to raw text in hand, before paste.
An association across these presses, not a fixed charge per minute, fitted only over the lengths observed. One speaker, hold mode, no AI polish; the lowest and highest tenth of waits dropped per length bin; different days and network load.
The line and its sample
| Measure | Value |
|---|---|
| Dictations over 60 s | 68 |
| Slope, s per minute | +1.54 |
| Intercept, s | 2.27 |
| R² | 0.88 |
| Longest, s | 431 |
Heavy files, few sessions
The audio had to get through my server too. A minute of uncompressed WAV weighed almost two megabytes, so a long dictation became a large upload before transcription could begin. At roughly thirteen minutes, raw audio reached the speech provider's file-size limit. That was another boundary to deal with, separate from the models' limits on audio duration and output text.
I load-tested the server by replaying one real dictation, about five minutes long, from increasing numbers of simulated users. One instance processed about eight requests at once; at fifteen users, 92% of requests failed. Doubling its memory moved the throughput ceiling by exactly zero. The constraints were the amount of request data it could handle per second and the number of requests its workers could process together.
I started by compressing the recording on the Mac, before sending it. Repeating the test with the same speech encoded as AAC, a much smaller file, reached forty simulated users without a failed request and more than quadrupled throughput. I stopped on the test's budget, not on the server: the new ceiling had not been found. Compression was a useful first fix, but the engine still received the whole recording only after I let go.
Compressing the same speech took one server instance from failing at 15 users to clean at 40.
One real 289.6 s dictation replayed by simulated users against one instance, 1 vCPU, 512 MiB.
| Measure | Raw WAV | Compressed AAC |
|---|---|---|
| File size | 9.27 MB | 907 KB, 24 kbps |
| Capacity ceiling, one instance | about 8 served (corrected by later calibration) | ceiling not reached at 40 offered |
| At 15 users | 92.4% failed (160 of 2,097 succeeded) | - |
| At 40 users | - | 0 failed of 240 |
| Whole 10-40 user ladder | - | 0 failed of 1,151 |
| Throughput, dictations / min | 55 | 254 |
| Doubling memory to 1 GiB | no throughput gain | - |
The compressed run stopped at the test's budget, not at a failure, so 40 users is not its ceiling. The app now ships 32 kbps, about 8 times smaller than raw; no load test exists after cutting.
Start before I finish
I was sending one whole recording after I stopped talking. Compression made it lighter, but none of its transcription happened while I spoke. Cutting offered a way to use that time: send a finished piece before the rest of the recording exists.
Each piece goes to the engine in an ordinary file request while I record the next one. The client keeps the returned text in recording order. Most transcription can finish before I let go; after release, the app waits for whatever is still outstanding, usually the final piece, then joins the text. The remaining work depends on those outstanding pieces rather than a whole recording that has only just been sent.
Transcribe each finished piece while I am still recording the next one.
The cutter's two-minute ceiling keeps each piece comfortably below the engines' length limits and keeps its audio upload small. Smaller bodies should also ease the byte-throughput constraint from the load tests, although I have no post-cutting capacity measurement. This addressed the request shape behind several problems at once. The next problem was the join: short files would not help much if splitting a word or a sentence made their transcripts harder to put back together.
First try: overlap and stitch
My earlier experiments used fixed-length pieces, so a cut could land in the middle of a word. I repeated a little audio on both sides of the cut to give the engine another chance at that speech. This returned two versions of the same passage. Joining them meant finding where the texts agreed and keeping the passage once.
The unfinished matcher searched too far back in the earlier text. Two common words in the wrong place convinced it that it had found the join, and it silently deleted 79 words of real speech, with no error raised. This time the missing words came from my own joining code. Restricting the search to the end of one piece and the beginning of the next repaired it. Across four long Ukrainian dictations, the bounded matcher showed a net word-count difference of four on 1,505 words, with no invented word sequences detected and every join reading coherently. It was a useful result from that tested set.
I also tried having a language model do the joining. Given the whole text, it produced visible duplication at the joins. Restricting it to the joins worked much better, but still deleted three real words and added paid model calls. The bounded algorithm did the job better. I had a workable way to join overlapping text, but every cut still created something to reconcile.
Four ways to join overlapping pieces, and what each did to the text
Earlier experiments on gpt-4o-mini-transcribe with overlapping pieces; each row has its own measure, so the rows share no common score.
| Method | Observed result | Conditions | Extra cost per audio hour | Outcome |
|---|---|---|---|---|
| Bounded matcher | net word count -4 of 1,505; no invented word sequences; 8 of 8 joins coherent | 4 longest Ukrainian dictations, 90 s pieces, 10 s overlap | +3.86% uk, +3.05% en | Worked, never shipped |
| Unbounded matcher | 79 real words deleted, silently | a false two-word match 46 words back; unfinished prototype | - | Fixed by bounding |
| Language model, whole text | visible duplication in 4 of 6 joins | gpt-4.1-mini | - | Dropped |
| Language model, joins only | 3 real words deleted | gpt-4.1-mini on the join regions | +10.71% uk, +11.19% en | Worse than bounded |
A net word-count difference is not four established omissions, nor an error score against a human reference. Costs compare with one whole-file call. The shipped cutter makes no overlap: none of these matchers ships.
Silence changes the question
The overlap was there because a fixed-length cut could land in the middle of a word. At a pause between phrases, there is no word to split: the pieces join end to end, in order, with no audio sent twice and nothing to match. My silence measurements showed such pauses are common: almost half the pauses between phrases lasted more than a second.
To hear a pause while I am still talking, the cutter uses the same kind of live detector as the recording indicator, with its own settings. It should call silence only when it is sure: a missed pause just makes a piece a little longer, while a false pause in the middle of a word cuts the word in two.
Cutting has a price of its own: each piece is transcribed without hearing the rest of the recording. I compared the text after cutting and trimming the pauses with the text of the same recordings sent whole, counting changed, added or missing words. There were 1.16 such differences per 100 words on average. Repeating the whole-file request produced 0.28 differences per 100 words. I did not check either transcript against what I actually said, so this comparison cannot tell which version got a word right. Cutting at pauses made the join simple; it did not give the engine back the rest of the recording.
How long should a piece be
My first rule waited forty seconds, then cut at the next pause. It had no ceiling, so nothing in the rule bounded a piece when a suitable pause failed to arrive. Later replays showed a forty-second start made many more joins with no measured gain, so the search moved to start after a minute, with a two-minute ceiling. These were search limits, not a schedule for cutting at exact intervals. Most short dictations stayed untouched: 88% of the tested archive went to the engine as one piece.
Then I looked at where the cuts actually landed. The typical piece lasted seventy-two seconds, not sixty, so I asked whether the floor could move down until the median cut fell on the minute. The study replayed 1,046 dictations and answered with a price. A forty-eight-second floor produced 44% more joins and 63% more final pieces shorter than fifteen seconds, for about 0.2 seconds less predicted median wait after release. That wait was modelled from the final pieces of recordings over two minutes, not newly measured through the app. That was a bad trade, and the one-minute start stayed.
Moving the floor later was not a clear win either. It reduced joins but increased the modelled worst wait as longer pieces became possible. The study recommended keeping both, with medium confidence. I had wanted smaller pieces, but it persuaded me: in my own use that balance is working well, and I am not planning to change it.
Moving the floor earlier buys many more joins for a fraction of a second; 60 / 120 stayed.
Replay of 1,046 dictations (11.6 h, one speaker). Joins and short final pieces: whole archive. Waits: modelled on the 55 recordings over 120 s.
| Floor / ceiling, s | Joins | Final pieces under 15 s | Modelled wait, median s | Modelled wait, worst s |
|---|---|---|---|---|
| 40 / 90 | 289 | 65 | 1.59 | 2.68 |
| 48 / 108 | 228 | 52 | 1.57 | 2.97 |
| 60 / 120, chosen | 158 | 32 | 1.78 | 2.97 |
| 75 / 135 | 108 | 22 | 1.78 | 3.80 |
| 90 / 150 | 83 | 17 | 1.73 | 4.10 |
Waits come from a model of the final piece (604 ms plus 23.7 ms per second of it), not from the app. The replay ran the live cutter before its voice look-back was added, and made no new engine calls.
Hearing a pause while you talk
The cutter has only the audio already recorded. Once a piece passes a minute, it waits for the first pause that begins after that point and lasts a full second. When the second is confirmed, it puts the boundary about halfway into the pause. The decision comes after the boundary, without needing to hear what I will say next.
At two minutes, it stops waiting and looks back over that piece's search window. First it chooses the longest pause of at least half a second. If the gate found none, the trimmer's voice detector gets one look at the recorded audio, looking for a non-speech stretch of at least a fifth of a second in the same window. I added that step after the archive replay produced two hard cuts in 158 cuts. In one recording, the live gate missed an earlier short gap; the voice scan moved the boundary from 119.9 seconds back to 82.1, out of the stretch of speech.
If even that scan finds no suitable boundary, the cutter uses the quietest point in the last ten seconds. That hard cut keeps the piece bounded, but can still divide speech. The gate can miss a real pause: in the unusually long waits I examined, sound above its silence line or a short burst inside a pause often stopped it accepting silence. Waiting longer was not always evidence that I had spoken without stopping.
Where the cutter cut one real two-minute dictation
Each piece's search window opens 60 s after its own start. The cutter cuts about 0.51 s into a pause, once it has confirmed the pause lasts a full second.
- 1cut at 124.99 s
- 1piece 1 ceiling, if no pause is found
- 2cut at 124.99 s
- 1piece 1 ceiling, if no pause is found
- 2cut at 124.99 s
- 3unused: piece 1 already cut
Both cuts found a full-second pause soon after their window opened, so neither piece reached its two-minute ceiling. Piece 2's window is counted from the first cut, not from a fixed two-minute grid.
The cuts, in seconds
| Piece | Window opens | Cut |
|---|---|---|
| 1 | 60 | 62.27 |
| 2 | 122.27 | 124.99 |
| last | - | release, 126.4 |
Fast, whatever the length
The change shows up where I wanted it: after I let go of the key. The chart plots my hold-to-talk recordings against the wait for raw text. Below a minute, both clouds still climb because nothing has been cut. Beyond a minute, whole-file recordings added about 1.7 seconds of wait per extra minute spoken on average across engines, and about 1.5 seconds on gpt-transcribe alone. Recordings sent in pieces stayed roughly flat over the lengths observed.
Wait after release against recorded length, before and after cutting
Showing 426 of 3,077 retained dictations: up to 16 actual recordings per engine, era and length bin. Dot density is not recording frequency.
- 1cutting starts after a minute
- 1cutting starts after a minute
- 1cutting starts after a minute
Each cloud mixes engines (whole file mostly gpt-4o-mini, pieces mostly Soniox and ElevenLabs); per-engine slopes below. Fits: all retained recordings over 60 s per era, over observed lengths. Retained: lowest and highest tenth of waits dropped per era, engine, length bin; AI polish, paste excluded. One speaker, room and language; hold mode; different weeks; network and provider load uncontrolled.
The counts and the lines behind the picture
| Era, engine | Retained | Shown | Over 60 s | Over 60 s, s/min | Longest, s |
|---|---|---|---|---|---|
| Whole file, all engines | 1,902 | 175 | 371 | +1.67 | 431 |
| In pieces, all engines | 1,175 | 251 | 280 | -0.11 | 608 |
| Whole file, gpt-4o-mini-transcribe | 1,597 | 91 | 303 | +1.71 | 373 |
| Whole file, gpt-transcribe | 305 | 84 | 68 | +1.54 | 431 |
| In pieces, gpt-transcribe | 171 | 76 | 35 | -0.01 | 248 |
| In pieces, Soniox | 680 | 90 | 170 | -0.12 | 608 |
| In pieces, ElevenLabs | 324 | 85 | 75 | -0.03 | 425 |
The two eras have different engine mixes: most earlier recordings used gpt-4o-mini-transcribe, while most later ones used Soniox or ElevenLabs. The shared gpt-transcribe comparison shows the same turn, and each cut engine has a near-flat line of its own above a minute. Before cutting, my five-to-ten-minute whole-file dictations across engines waited a median of about ten seconds, the worst nearly half a minute.
The cuts also have a live check in my production and development logs: 323 of 331 landed at a full pause, the remainder at shorter ones, and none required a hard cut. My longest ordinary dictation so far was a thought I dictated for a little over ten minutes. Soniox handled it in eight pieces, and the raw text was in hand 4.4 seconds after release. Dictations that long are still rare for me.
- raw text after release, a ten-minute dictation in eight pieces
- 4.4 s
- cuts that landed at a full-second pause
- 323 of 331
- hard cuts
- 0 of 331
4.4 s: one Soniox dictation of a little over ten minutes, before paste. Cuts: 234 cut dictations in my production and development logs; observations, not a guarantee.
An hour on the bench
My real dictations had reached about ten minutes, so testing an hour meant assembling one. I joined 124 real recordings into just under an hour of audio, with room noise between them, and ran it through the desktop's recording code with gpt-transcribe and a local API. In two raw-text runs, it made 44 pieces. At speaking speed, only a short final piece remained after release; when I fed the audio ten times faster, an earlier piece was still in flight too. Raw text was ready in 0.9-3.2 seconds across the runs, depending on the outstanding pieces. A similarly long file took 92 seconds in the separate whole-file provider test. Most of the piecewise pipeline's work had happened while the recording was still going.
That timing does not certify a complete transcript. The test used an earlier trimmer, and 122 of the 124 clips came back in the text. It has not been repeated on the current build. Both missing clips were the same quiet whisper; one of them the engine loses even untrimmed. The silence-trimming article covers the trimmer's part. These raw-text timings exclude the production network hop and paste. AI formatting of the whole text adds its own wait, covered below.
What the pieces open up
While I am still speaking, the app already holds text from the pieces that have returned, in order, on my Mac. Today the pieces stay invisible; the user still gets one dictation. Having that text early makes room for a future interface to show progress through a long recording, based on work already finished.
These are ordinary file requests, rather than a realtime transcription session. In my comparison, realtime cost more on Soniox, ElevenLabs and Deepgram, though not on gpt-transcribe; Soniox's realtime mode also bills the whole session, silence included, so trimming pauses would save nothing there. Cutting let me keep the file engines and finish most of their work before release. In my longer dictations, I can keep following the thought, then let go and have the raw text in a few seconds.
A few practical questions
How long can I dictate?
While you hold the key, Shepit records for up to an hour. In toggle mode, one press starts recording and the next stops it; the default cap is ten minutes, adjustable in Settings to an hour. Cutting keeps the engine requests short, but these recording caps still apply.
Will I notice the pieces?
No. The app cuts, sends and joins them in order. You still get one dictation when you finish; it does not show provisional text while you speak.
What happens if one piece fails?
Retry sends only pieces that have no text yet and keeps those already returned. It repairs absent responses; it does not check returned text for wrong or missing words.
Does AI formatting get faster too?
No. Formatting works on the full text after the last piece returns. In the assembled-hour test, the formatted result was ready about 80 seconds after release, compared with 0.9-3.2 seconds for raw text. That full-text pass remains after recording.
Does cutting depend on the language I speak?
The cutter's rule listens for pauses in the sound, rather than recognizing words. My live-use wait and cut measurements come from Ukrainian dictations by one speaker: me. They do not establish how well it works across languages.

