Why pay for silence?
Quick answer
About 10% less audio went to transcription across 1,925 real dictations. Shepit shortens pauses longer than a second to one second. In a small test with checked transcripts, that setting did not increase word errors.
I was building a small thing: the capsule that appears on screen while you hold Shepit's dictation key. Its bars rise and fall with your voice. When you stop talking, it narrows. Release the key, and Shepit sends the recording for transcription.

The pill that saw the silence
The first version never narrowed in my room. I could stop speaking, but the ventilation's low hum stayed above the level I had chosen for silence. Teaching the pill to learn the room's background level brought the opposite problem: now it could narrow in the middle of a sentence. Measuring where my voice stood out from the rumble fixed it. I tried to trick it by speaking deliberately quietly, and the pill stayed open.
Once it worked, the pill made another part of dictation visible: the thinking. I lose the thread, wait for the next thought, and the bars settle down while the recording keeps going. Sometimes you know what you mean before you know how to say it. I want room for that.
The transcription engine received that thinking time too. On a duration-billed engine, a second without speech costs the same as a second with it. Even the model I used then, which counted audio tokens rather than seconds, charged for the extra audio. A smaller pill did nothing to that bill.
If we can tell where the silence is, why are we paying to transcribe it?
First, measure
A quick script judged silence by loudness alone. Across 536 recordings, about 4.7 hours of audio, it said I could remove 17.5%. I had not checked whether the words survived. One recording exposed the problem: the rule would remove three quarters of a roughly 90-second clip from a loud room. A large saving can mean a bad detector.
So I measured again with a detector that looked for a voice rather than for quiet. In 20-millisecond slices, it checked for energy in speech frequencies jumping above the room's steady background, and a repeating pattern at a human voice's pitch. Both checks had to pass to find a voice; the detector then kept nearby unvoiced sounds too.
Two signs of speech
An example recording, measured by the current detector
- 1Pitch check passes for 2 s; energy stays below 10 dB
- 1Pitch check passes for 2 s; energy stays below 10 dB
- 110 dB: may be speech
- 2Pitch check passes for 2 s; energy stays below 10 dB
- 3pitch-check threshold
The first measurement used these energy and pitch checks too. The same recording appears in the trimming example later. This is a current-detector illustration, not a reconstruction of the earlier measurement.
In that first measurement, the detector classified almost a quarter of just over ten hours of dictation as non-speech. Little of it sat before the first word or after the last. Most of it was between phrases: the thinking time I had been watching in the pill.
Pauses inside 1,046 dictations, by length
1,046 dictations, 10.05 h; non-speech: 20.3% internal pauses, 2.3% before speech, 1.5% after. Earlier detector version; colours mark what the shipped rule would shorten.
1,000 dictations were mine. Median pause: 0.94 s. Pauses over a second were 47.2% of pauses and occupied 15.8% of all audio, recomputed from per-recording gap lists. Pause count is not audio duration; no bar measures safe saving.
The pauses behind the picture
| Pause length | Pauses |
|---|---|
| Under 0.5 s | 1,082 |
| 0.5–1 s | 1,710 |
| 1–2 s | 1,514 |
| 2–5 s | 820 |
| Over 5 s | 160 |
A pause between thoughts also separates phrases. Shortening it moves those phrases closer together. The measurement had found plenty to cut; it had not shown how short those gaps could become before the transcript changed.
I got greedy: half a second
I started with a greedy rule for pauses inside the recording: every pause longer than half a second becomes half a second. On the same ten hours of dictation, that would remove about 16% of the audio. There would still be a gap between phrases. I had priced the cut without yet transcribing the shortened audio.
One of the engines I tested, Soniox, made that trade look less attractive. On 76 dictations with checked transcripts, half-second pauses took its score from 0.57 errors per 100 words to 0.95. Even one second nudged the errors up. A second and a half matched the untrimmed score in that test, so I moved the plan to that setting. It still offered about 9% less audio.
The pill's ear could not do the cutting
The code that split long recordings already used the pill's live signal, the one that noticed when I stopped talking. Borrowing it looked simpler than building another detector.
But the pill was built to keep a display steady. A pill that blinks shut mid-word looks broken; one that stays open a little too long costs nothing visible. So I had made it reluctant to close. For trimming, that caution left pauses untouched. Under the same 1.5-second rule, it would remove less than a third as much audio as the dedicated detector.
Audio each detector would remove
Offline replay of the same 605 recordings; internal pauses shortened to 1.5 s
This compares removed audio under one rule, not accuracy. It is not the shipped one-second setting.
Same recordings, same pause length
| Detector | Audio removed, % |
|---|---|
| Pill gate | 2.67 |
| Dedicated detector | 8.58 |
It could also make the opposite mistake. In a test, borrowing the pill's background-noise estimate removed 70% of a very quiet dictation. When I started it without a learned background level, it cut almost the whole recording, and the transcript came back empty. A display mistake had become missing audio.
So the trimmer got its own detector: the voice-aware approach from the measurement, run on each finished piece of audio. With the piece complete, it could look ahead before deciding where to cut.
Back to one second
With the separate detector in place, I repeated the checks on gpt-transcribe, the engine Shepit was now using. On the 76 dictations with checked transcripts, half-second pauses scored 2.08 errors per 100 words against 1.93 untrimmed: four more errors in 2,643 words. That was too small a difference to settle the choice on its own.
The long dictations gave me a different check. Shepit splits them into pieces, so on 207 recordings longer than a minute I compared the text from the trimmed pieces with two whole-file sends of the same audio.
The half-second version added missing-word warnings
207 dictations over 60 s; compared with both whole-file sends, not a human reference
| Pipeline | Recordings with a 3+ word run missing | What the comparison found |
|---|---|---|
| Split, no trim | 3 | Baseline differences from sending in pieces. |
| Split, 1.0 s pauses | 5 | Mostly self-corrections or repeats dropped; none was at a join, where trimmed audio is stitched together. |
| Split, 0.5 s pauses | 7 | Added a seven-word run and a whole missing sentence. |
These counts come from a separate test from the 76 scored recordings.
Splitting alone already produced missing-word warnings; shortening the pauses added more. Not every warning was a real error: there was no human reference, and some were self-corrections or repeats the model dropped. The added seven-word run and missing sentence were warnings I could not dismiss for a better saving.
There was also a simpler option: trim only before the first word and after the last, leaving every gap between phrases intact. That would avoid internal joins, but it would leave most of the potential saving untouched.
In the 76-recording test, one second and a second and a half were harder to separate. One second came out a single error ahead. Then I sent the exact same one-and-a-half-second files again, and the count moved by two. The same bytes moved the score more than the setting did.
In this sample, quality could not choose between them, but the durations could. On a separate 11.6-hour archive of real dictations, one second removed 11.5% of the audio against 8.9% at a second and a half. I chose one second, knowing it nearly doubled the joins in the 76-recording test. That was a risk this small sample could not settle.
The detector ate words
I stitched 124 real recordings into one nearly hour-long file for an artificial stress test. One was the same very quiet dictation the pill's noise estimate had cut by 70%. Trimmed alone with the new detector, it kept 77% of its audio. Inside the assembled file, it disappeared completely. Sent untrimmed, it came back.
A quiet clip disappeared inside a louder recording
Audio retained before the fix; one clip alone and in a spliced 124-clip, 59:52 stress test
Found before release.
The retained audio
| Test condition | Quiet clip retained, % |
|---|---|
| Trimmed alone | 77 |
| Inside the loud file | 0.0 |
The words had not changed. Their loudness relative to everything around them had. The quiet dictation sat about 35 dB below its neighbours, and the detector ignored any sound more than 30 dB below the loudest moment of its piece. Stitching recordings made at different gains had created that contrast. That cutoff was already part of the first measurement. The artificial test exposed its bad assumption, so I removed it.
In an earlier check, four spoken words had been cut across three pieces. All three went through the special handling for very noisy audio. There I stopped trimming: those pieces now go whole. Both fixes were in place before release.
On a separate replay of 937 dictations, the fixes reduced the removed share from about 11.5% to 10.4%. I had gone looking for a cheaper upload and ended up deliberately keeping more audio.
What ships: a pause longer than a second becomes one second, with 0.3 seconds kept after the phrase and 0.7 before the next. The start and end keep 0.2 seconds: there is no neighbouring phrase there for a pause to separate.
Word-error scores stayed close across the tested settings
gpt-transcribe; 76 dictations, 2,643 reference words; one scored run per setting, plus one repeat
- 11.5 s, same files again: same bytes
- 21.0 s: ships
- 11.5 s, same files again: same bytes
The report treats differences under about ten errors as unresolved on this sample. One repeat; not a confidence interval.
The scored results
| Setting | Errors | Errors per 100 words |
|---|---|---|
| Untrimmed | 51 | 1.93 |
| 1.5 s | 49 | 1.85 |
| 1.5 s, same files again | 51 | 1.93 |
| 1.0 s | 48 | 1.82 |
| 0.5 s | 55 | 2.08 |
Shortening pauses leaves space between phrases
One 24-second dictation, 4.7 seconds removed; example, not the average saving
- 1Leading silence cut to 0.2 s.
- 2A 0.62 s pause, under a second: left whole.
- 3A 3.0 s pause becomes 1.0 s: 0.3 s after, 0.7 s before.
- 4Trailing silence cut to 0.2 s; 10 ms fades at joins.
- 1Leading silence cut to 0.2 s.
- 2A 0.62 s pause, under a second: left whole.
- 3A 3.0 s pause becomes 1.0 s: 0.3 s after, 0.7 s before.
- 4Trailing silence cut to 0.2 s; 10 ms fades at joins.
- 1Leading silence cut to 0.2 s.
- 2A 0.62 s pause, under a second: left whole.
- 3A 3.0 s pause becomes 1.0 s: 0.3 s after, 0.7 s before.
- 4Trailing silence cut to 0.2 s; 10 ms fades at joins.
This example illustrates the rule, not a quality test.
The removed intervals
| Removed interval, s | Removed, s |
|---|---|
| 0.00–0.38 | 0.38 |
| 1.24–1.34 | 0.10 |
| 8.54–10.54 | 2.00 |
| 19.96–21.62 | 1.66 |
| 23.44–24.00 | 0.56 |
| Total | 4.70 |
What it saves in real use
The server records both durations for each dictation: the original recording and the audio sent for transcription. The difference measures how much audio the trimmer removed in real use.
- recorded audio removed
- 10.55 %
- dictations checked
- 1,925
- audio recorded
- 20.68 h
- audio sent
- 18.50 h
Nine accounts; my dictations supplied 84% of recorded time. Nothing removed in 15.7% of dictations. Durations do not measure transcription accuracy.
Most of those hours were mine. I dictate long pieces and leave thinking pauses. The other eight accounts saved a smaller share, but they also dictated shorter pieces. This comparison cannot tell me whether the difference came from the speaker or the length of the dictation. About 10% is what this mix of recordings gave, not a forecast for everyone who uses Shepit.
Audio removed in real use, by length and speaker
Share of recorded duration removed within each group; 1,925 dictations, 20.68 recorded hours
Length rows exclude the 12 dictations over five minutes; speaker rows include all lengths. The other accounts dictated shorter pieces; this comparison does not isolate length from speaker.
The groups behind the saving
| Group | Audio removed, % |
|---|---|
| Under 15 s | 7.2 |
| 15–60 s | 9.5 |
| 1–5 min | 12.1 |
| My dictations | 11.74 |
| Eight other accounts | 4.34 |
The pill could not shrink the bill, but shorter uploads can. At gpt-transcribe's $0.27 per audio hour, the measured reduction works out to roughly $28 per 1,000 recorded hours. That is a calculation from the durations, not a saving I have checked against invoices.
I still record the whole dictation while I speak. The cuts happen afterwards, before the upload. I can lose the thread and wait for the next thought just as before; the transcription engine simply receives less of that waiting.
A few practical questions
Why shorten pauses instead of deleting them?
A pause separates phrases, so I keep a gap. In a separate experiment that split recordings at half-second pauses, sentence breaks changed at 87 of 171 split points compared with whole-file text. That was a splitting test, not a trimming test, and it did not establish which punctuation was right. Keeping a gap is a design choice; the trimming tests helped me choose its length.
Does trimming make dictation slower?
The test found no clear slowdown for recordings under a minute. In a 700-request latency test with three repeats, sub-minute dictations changed by only 0.02-0.03 seconds at the median, within run-to-run variation. For recordings of 60-120 seconds, the measured tail latency improved by about 0.27 seconds. The trimming pass itself took 94 ms per 120-second piece in one local benchmark on an M1 Pro. These are results for that setup, not a guarantee for every machine.
What happens when a piece is too noisy to trim reliably?
If its loudest and quietest levels are less than 20 dB apart, Shepit sends the piece whole. That condition replaced special handling for very noisy audio that cut four spoken words across three pieces in tests. There is no trimming saving on that piece.
How could I test silence trimming on my own transcription engine?
Compare original and trimmed recordings against checked transcripts. Repeat identical sends too: if the same bytes change the score, the engine is adding variation to the comparison. For long recordings, compare the piecewise result with whole-file sends, while remembering that model output is not a human reference. Count joins as well as seconds removed. Try an artificial stress test that puts very quiet recordings next to loud ones; it can expose a cutoff tied to the loudest sound. These are checks I used in this investigation.

