Why pay for silence?

Quick answer

About 10% less audio went to transcription across 1,925 real dictations. Shepit shortens pauses longer than a second to one second. In a small test with checked transcripts, that setting did not increase word errors.

Published
Reading time
11 min

I was building a small thing: the capsule that appears on screen while you hold Shepit's dictation key. Its bars rise and fall with your voice. When you stop talking, it narrows. Release the key, and Shepit sends the recording for transcription.

Two Shepit dictation pills at their real proportions: on the left, expanded with waveform bars while speaking; on the right, narrowed to four short bars during a pause.
The dictation pill expands with speech and narrows during pauses. Its display follows the recording; it does not cut the audio.

The pill that saw the silence

The first version never narrowed in my room. I could stop speaking, but the ventilation's low hum stayed above the level I had chosen for silence. Teaching the pill to learn the room's background level brought the opposite problem: now it could narrow in the middle of a sentence. Measuring where my voice stood out from the rumble fixed it. I tried to trick it by speaking deliberately quietly, and the pill stayed open.

Once it worked, the pill made another part of dictation visible: the thinking. I lose the thread, wait for the next thought, and the bars settle down while the recording keeps going. Sometimes you know what you mean before you know how to say it. I want room for that.

The transcription engine received that thinking time too. On a duration-billed engine, a second without speech costs the same as a second with it. Even the model I used then, which counted audio tokens rather than seconds, charged for the extra audio. A smaller pill did nothing to that bill.

If we can tell where the silence is, why are we paying to transcribe it?

First, measure

A quick script judged silence by loudness alone. Across 536 recordings, about 4.7 hours of audio, it said I could remove 17.5%. I had not checked whether the words survived. One recording exposed the problem: the rule would remove three quarters of a roughly 90-second clip from a loud room. A large saving can mean a bad detector.

So I measured again with a detector that looked for a voice rather than for quiet. In 20-millisecond slices, it checked for energy in speech frequencies jumping above the room's steady background, and a repeating pattern at a human voice's pitch. Both checks had to pass to find a voice; the detector then kept nearby unvoiced sounds too.

Two signs of speech

An example recording, measured by the current detector

detected speechthe recordingenergy in speech frequencies, dB above the room's background10 dB: may be speechhow strongly the sound repeats at a voice's pitchpitch-check threshold1detected speech, including unvoiced edges681012 s
  1. 1Pitch check passes for 2 s; energy stays below 10 dB
detected speechthe recordingenergy in speech frequencies, dB above the room's background10 dB: may be speechhow strongly the sound repeats at a voice's pitchpitch-check threshold1detected speech, including unvoiced edges681012 s
  1. 1Pitch check passes for 2 s; energy stays below 10 dB
detected speechthe recordingenergy in speech frequencies, dB above theroom's background1how strongly the sound repeats at avoice's pitch23detected speech, including unvoiced edges681012 s
  1. 110 dB: may be speech
  2. 2Pitch check passes for 2 s; energy stays below 10 dB
  3. 3pitch-check threshold

The first measurement used these energy and pitch checks too. The same recording appears in the trimming example later. This is a current-detector illustration, not a reconstruction of the earlier measurement.

In that first measurement, the detector classified almost a quarter of just over ten hours of dictation as non-speech. Little of it sat before the first word or after the last. Most of it was between phrases: the thinking time I had been watching in the pill.

Pauses inside 1,046 dictations, by length

1,046 dictations, 10.05 h; non-speech: 20.3% internal pauses, 2.3% before speech, 1.5% after. Earlier detector version; colours mark what the shipped rule would shorten.

05001,0001,5002,000 pausesLeft wholeUnder 0.5 s1,082 pauses0.5–1 s1,710 pausesShortened1–2 s1,514 pauses2–5 s820 pausesOver 5 s160 pauses
05001,0002,000 pausesLeft wholeUnder 0.5 s1,082 pauses0.5–1 s1,710 pausesShortened1–2 s1,514 pauses2–5 s820 pausesOver 5 s160 pauses

1,000 dictations were mine. Median pause: 0.94 s. Pauses over a second were 47.2% of pauses and occupied 15.8% of all audio, recomputed from per-recording gap lists. Pause count is not audio duration; no bar measures safe saving.

The pauses behind the picture
Pause lengthPauses
Under 0.5 s1,082
0.5–1 s1,710
1–2 s1,514
2–5 s820
Over 5 s160

A pause between thoughts also separates phrases. Shortening it moves those phrases closer together. The measurement had found plenty to cut; it had not shown how short those gaps could become before the transcript changed.

I got greedy: half a second

I started with a greedy rule for pauses inside the recording: every pause longer than half a second becomes half a second. On the same ten hours of dictation, that would remove about 16% of the audio. There would still be a gap between phrases. I had priced the cut without yet transcribing the shortened audio.

One of the engines I tested, Soniox, made that trade look less attractive. On 76 dictations with checked transcripts, half-second pauses took its score from 0.57 errors per 100 words to 0.95. Even one second nudged the errors up. A second and a half matched the untrimmed score in that test, so I moved the plan to that setting. It still offered about 9% less audio.

The pill's ear could not do the cutting

The code that split long recordings already used the pill's live signal, the one that noticed when I stopped talking. Borrowing it looked simpler than building another detector.

But the pill was built to keep a display steady. A pill that blinks shut mid-word looks broken; one that stays open a little too long costs nothing visible. So I had made it reluctant to close. For trimming, that caution left pauses untouched. Under the same 1.5-second rule, it would remove less than a third as much audio as the dedicated detector.

Audio each detector would remove

Offline replay of the same 605 recordings; internal pauses shortened to 1.5 s

0246810%Pill gate2.67%Dedicated detector8.58%
0246810%Pill gate2.67%Dedicated detector8.58%

This compares removed audio under one rule, not accuracy. It is not the shipped one-second setting.

Same recordings, same pause length
DetectorAudio removed, %
Pill gate2.67
Dedicated detector8.58

It could also make the opposite mistake. In a test, borrowing the pill's background-noise estimate removed 70% of a very quiet dictation. When I started it without a learned background level, it cut almost the whole recording, and the transcript came back empty. A display mistake had become missing audio.

So the trimmer got its own detector: the voice-aware approach from the measurement, run on each finished piece of audio. With the piece complete, it could look ahead before deciding where to cut.

Back to one second

With the separate detector in place, I repeated the checks on gpt-transcribe, the engine Shepit was now using. On the 76 dictations with checked transcripts, half-second pauses scored 2.08 errors per 100 words against 1.93 untrimmed: four more errors in 2,643 words. That was too small a difference to settle the choice on its own.

The long dictations gave me a different check. Shepit splits them into pieces, so on 207 recordings longer than a minute I compared the text from the trimmed pieces with two whole-file sends of the same audio.

The half-second version added missing-word warnings

207 dictations over 60 s; compared with both whole-file sends, not a human reference

PipelineRecordings with a 3+ word run missingWhat the comparison found
Split, no trim3Baseline differences from sending in pieces.
Split, 1.0 s pauses5Mostly self-corrections or repeats dropped; none was at a join, where trimmed audio is stitched together.
Split, 0.5 s pauses7Added a seven-word run and a whole missing sentence.

These counts come from a separate test from the 76 scored recordings.

Splitting alone already produced missing-word warnings; shortening the pauses added more. Not every warning was a real error: there was no human reference, and some were self-corrections or repeats the model dropped. The added seven-word run and missing sentence were warnings I could not dismiss for a better saving.

There was also a simpler option: trim only before the first word and after the last, leaving every gap between phrases intact. That would avoid internal joins, but it would leave most of the potential saving untouched.

In the 76-recording test, one second and a second and a half were harder to separate. One second came out a single error ahead. Then I sent the exact same one-and-a-half-second files again, and the count moved by two. The same bytes moved the score more than the setting did.

In this sample, quality could not choose between them, but the durations could. On a separate 11.6-hour archive of real dictations, one second removed 11.5% of the audio against 8.9% at a second and a half. I chose one second, knowing it nearly doubled the joins in the 76-recording test. That was a risk this small sample could not settle.

The detector ate words

I stitched 124 real recordings into one nearly hour-long file for an artificial stress test. One was the same very quiet dictation the pill's noise estimate had cut by 70%. Trimmed alone with the new detector, it kept 77% of its audio. Inside the assembled file, it disappeared completely. Sent untrimmed, it came back.

A quiet clip disappeared inside a louder recording

Audio retained before the fix; one clip alone and in a spliced 124-clip, 59:52 stress test

Trimmed alone → Inside the loud file0255075100%Quiet counting rhyme77% → 0%
Trimmed alone → Inside the loud file0255075100%Quiet counting rhyme77% → 0%

Found before release.

The retained audio
Test conditionQuiet clip retained, %
Trimmed alone77
Inside the loud file0.0

The words had not changed. Their loudness relative to everything around them had. The quiet dictation sat about 35 dB below its neighbours, and the detector ignored any sound more than 30 dB below the loudest moment of its piece. Stitching recordings made at different gains had created that contrast. That cutoff was already part of the first measurement. The artificial test exposed its bad assumption, so I removed it.

In an earlier check, four spoken words had been cut across three pieces. All three went through the special handling for very noisy audio. There I stopped trimming: those pieces now go whole. Both fixes were in place before release.

On a separate replay of 937 dictations, the fixes reduced the removed share from about 11.5% to 10.4%. I had gone looking for a cheaper upload and ended up deliberately keeping more audio.

What ships: a pause longer than a second becomes one second, with 0.3 seconds kept after the phrase and 0.7 before the next. The start and end keep 0.2 seconds: there is no neighbouring phrase there for a pause to separate.

Word-error scores stayed close across the tested settings

gpt-transcribe; 76 dictations, 2,643 reference words; one scored run per setting, plus one repeat

errors per 100 wordsthe same files sent again11.522.5Untrimmed1.931.5 s1.851.5 s, same files again1.93 · same bytes1.0 s1.82 · ships0.5 s2.08
errors per 100 wordsthe same files sent again11.522.5Untrimmed1.931.5 s1.851.5 s, same files again1.9311.0 s1.8220.5 s2.08
  1. 11.5 s, same files again: same bytes
  2. 21.0 s: ships
errors per 100 wordsthe same files sent again11.522.5Untrimmed1.931.5 s1.851.5 s, same files again1.9311.0 s1.82 · ships0.5 s2.08
  1. 11.5 s, same files again: same bytes

The report treats differences under about ten errors as unresolved on this sample. One repeat; not a confidence interval.

The scored results
SettingErrorsErrors per 100 words
Untrimmed511.93
1.5 s491.85
1.5 s, same files again511.93
1.0 s481.82
0.5 s552.08

Shortening pauses leaves space between phrases

One 24-second dictation, 4.7 seconds removed; example, not the average saving

speechnot speechremovedbefore: 24.0 s123406121824 sAfter: 19.3 s, about 20% shorter06121819.3 s
  1. 1Leading silence cut to 0.2 s.
  2. 2A 0.62 s pause, under a second: left whole.
  3. 3A 3.0 s pause becomes 1.0 s: 0.3 s after, 0.7 s before.
  4. 4Trailing silence cut to 0.2 s; 10 ms fades at joins.
speechnot speechremovedbefore: 24.0 s123406121824 sAfter: 19.3 s, about 20% shorter061219.3 s
  1. 1Leading silence cut to 0.2 s.
  2. 2A 0.62 s pause, under a second: left whole.
  3. 3A 3.0 s pause becomes 1.0 s: 0.3 s after, 0.7 s before.
  4. 4Trailing silence cut to 0.2 s; 10 ms fades at joins.
speechnot speechremovedbefore: 24.0 s123406121824 sAfter: 19.3 s, about 20% shorter061219.3 s
  1. 1Leading silence cut to 0.2 s.
  2. 2A 0.62 s pause, under a second: left whole.
  3. 3A 3.0 s pause becomes 1.0 s: 0.3 s after, 0.7 s before.
  4. 4Trailing silence cut to 0.2 s; 10 ms fades at joins.

This example illustrates the rule, not a quality test.

The removed intervals
Removed interval, sRemoved, s
0.00–0.380.38
1.24–1.340.10
8.54–10.542.00
19.96–21.621.66
23.44–24.000.56
Total4.70

What it saves in real use

The server records both durations for each dictation: the original recording and the audio sent for transcription. The difference measures how much audio the trimmer removed in real use.

recorded audio removed
10.55 %
dictations checked
1,925
audio recorded
20.68 h
audio sent
18.50 h

Nine accounts; my dictations supplied 84% of recorded time. Nothing removed in 15.7% of dictations. Durations do not measure transcription accuracy.

Most of those hours were mine. I dictate long pieces and leave thinking pauses. The other eight accounts saved a smaller share, but they also dictated shorter pieces. This comparison cannot tell me whether the difference came from the speaker or the length of the dictation. About 10% is what this mix of recordings gave, not a forecast for everyone who uses Shepit.

Audio removed in real use, by length and speaker

Share of recorded duration removed within each group; 1,925 dictations, 20.68 recorded hours

051015%By lengthUnder 15 s7.2%15–60 s9.5%1–5 min12.1%By speakerMy dictations11.74%Eight other accounts4.34%
051015%By lengthUnder 15 s7.2%15–60 s9.5%1–5 min12.1%By speakerMy dictations11.74%Eight other accounts4.34%

Length rows exclude the 12 dictations over five minutes; speaker rows include all lengths. The other accounts dictated shorter pieces; this comparison does not isolate length from speaker.

The groups behind the saving
GroupAudio removed, %
Under 15 s7.2
15–60 s9.5
1–5 min12.1
My dictations11.74
Eight other accounts4.34

The pill could not shrink the bill, but shorter uploads can. At gpt-transcribe's $0.27 per audio hour, the measured reduction works out to roughly $28 per 1,000 recorded hours. That is a calculation from the durations, not a saving I have checked against invoices.

I still record the whole dictation while I speak. The cuts happen afterwards, before the upload. I can lose the thread and wait for the next thought just as before; the transcription engine simply receives less of that waiting.

Written by

Oleksij Rak

Co-founder of Rebbix, maker of Shepit

Engineer for 30 years. Started Shepit to fix his own dictation and builds it hands-on, from the speech pipeline to the tools behind this blog.

Published