sami7777 612c184317 Add silence-trim subcommand (alias st): collapse long pauses, v7-hardened
- sermon_clean/silence_trim.py: function-form silence_trim(audio, output, ...) -> SilenceTrimResult
  with full v7 guards (silencedetect nonzero-exit raise, min>keep validation, full-N
  lossless WAV complexity trial, >5% drift abort)
- sermon_clean/cli.py: cmd_silence_trim() + p_st parser registered before p_pipe
- Verified end-to-end on synthetic 20-cycle 9.00s WAV (drift 0.13%, complexity trial ran)
- Invalid-range rejections verified: min<=keep, min<=0, test_segments<=0 → rc=1
- Source: closed prior-tick pickup from _drafts/audio-silence-trim/ SKILL.md
2026-07-29 01:06:46 -07:00
2026-07-27 02:40:18 -07:00
2026-07-27 02:40:18 -07:00

sermon-clean

A one-shot CLI for sermon audio editing: find bad segments, cut them out, paste in ElevenLabs replacements. Built because doing this by hand every time is unbearable.

Why

Recording a sermon is fine. Post-production is not. The current workflow needs:

  1. Manually identify where the bad words are (the painful part — Whisper stalls on long audio, and eyeballing a waveform is imprecise)
  2. Hand-write ffmpeg trim commands for each bad window
  3. Render replacement clips via ElevenLabs
  4. Hand-write the ffmpeg concat command
  5. Manually upload to Dropbox

sermon-clean collapses steps 1-5 into one command (or a few, if you want to eyeball the audio first).

Install

# Requires ffmpeg in PATH (sudo apt install ffmpeg on Debian/Ubuntu)
pip install sermon-clean

# Optional: for ElevenLabs auto-rendering
export ELEVENLABS_API_KEY=...
export ELEVENLABS_VOICE_ID=IYUnpZr9CQfSylOsOOBo  # your cloned voice

# Optional: for auto-detect mode (transcribe + find bad words)
pip install "sermon-clean[auto]"

Usage

One-shot: explicit timestamps + replacement text

sermon-clean pipe sermon.ogg \
    --bad "21:38-21:42:the actual sentence you meant to say" \
    --bad "1450.3-1453.1:the corrected phrase" \
    --elevenlabs-text \
    -o sermon-fixed.ogg

Step-by-step workflow

# 1. Look at the audio — see its shape + find silence gaps
sermon-clean scan sermon.ogg --width 100

# 1b. Or: get a tabular index of silence runs (good for picking splice points)
sermon-clean si sermon.ogg --silence-threshold -35 --silence-min-duration 0.5

# 1c. Or: multiband waveform — N seconds per row, makes timestamp counting trivial
sermon-clean mb sermon.ogg --band-seconds 60 --width 80

# 1d. Or: silence stats — count/mean/longest silence. Use this to pick the right
# threshold for the next recording of the same speaker/room setup.
sermon-clean silence-stats sermon.ogg

# 1e. Or: auto-tune the silence threshold based on expected pause density.
# Useful when the same speaker records in different rooms and the silence
# profile changes week to week.
sermon-clean threshold-tune sermon.ogg --target-spm 4.0

# 1f. Or: pre-flight normalization — bring the recording to broadcast-standard
# loudness before doing any other processing.
sermon-clean normalize sermon.ogg -o sermon-normalized.ogg

# 1g. Or: light denoise (FFT-based) if the recording has hiss / mic preamp noise.
sermon-clean denoise sermon.ogg -o sermon-denoised.ogg

# 2. Or: extract overlapping slices for manual review
sermon-clean slices sermon.ogg --output-dir ./slices \
    --slice-seconds 5 --overlap-seconds 1
# Listen to ./slices/slice_000_0-00-0-05.ogg in your audio player.
# Filename embeds start/end timestamps.

# 3. Edit the JSON to mark bad segments + replacement text:
# segs.json:
# [
#   {"start": 1298.0, "end": 1302.0, "reason": "misspoke", "replacement_text": "the actual sentence"},
#   {"start": 1450.3, "end": 1453.1, "reason": "misspoke", "replacement_text": "the corrected phrase"}
# ]

# 4. Trim the original around the bad windows
sermon-clean cut sermon.ogg --segments-file segs.json

# 5. Render replacements (if you haven't already) and splice
sermon-clean paste sermon.ogg --segments-file segs.json \
    --replacements "replacements/*.mp3" \
    -o sermon-fixed.ogg

Auto-detect mode (optional, requires faster-whisper)

sermon-clean auto sermon.ogg --bad-words "fuck,shit,damn" --output segs.json
# transcribes the audio, finds timestamps for any of the bad words,
# prints a suggested JSON file you can edit before splicing

Subcommands

Command Purpose
find Show audio metadata + silence gaps (no transcription)
scan ASCII waveform + silence marks (no transcription)
silence-index (si) Tabular list of silence runs with timestamps + position bar
multiband (mb) Multi-row ASCII waveform with band-start labels for timestamp counting
silence-stats (ss) Quantitative summary: count, mean, median, longest silence + density per minute
threshold-tune (tt) Auto-pick the silence threshold that matches expected pause density
normalize (norm) Apply EBU R128 loudness normalization (target LUFS, true peak, LRA)
denoise Apply light FFT-based noise reduction (afftdn) for hiss / mic preamp noise
slices Extract overlapping audio chunks for manual review
cut Trim the original around bad windows (no splice)
paste Splice pre-rendered replacements into the trimmed original
auto Transcribe + find bad-word timestamps
pipe Run cut + paste in one command
batch Run any of the above across many files

Batch processing

Apply the same operation to many files at once:

# Normalize every sermon from this Sunday
sermon-clean batch normalize 'sermons/*.ogg' --output-dir fixed/ --suffix=-normalized

# Denoise a batch of older recordings
sermon-clean batch denoise 'archive/*.wav' --output-dir fixed/ --suffix=-dn

# Silence stats for every recording (no output files — runs the stats print)
sermon-clean batch silence-stats 'sermons/*.ogg' --output-dir stats/

Failures are collected, not raised: if one file is corrupt, the rest still process.

How it works

  • Findffmpeg silencedetect for natural breath pauses, plus your explicit timestamps
  • Cut → trim the original into N+1 good segments using ffmpeg, re-encoding to the source codec (no wav intermediate — bitrate-matched)
  • Paste → ElevenLabs renders + ffmpeg concat with bit-exact -c copy (no audible clicks at splice boundaries)
  • Verify → ffprobe the output duration against expected; if it drifted >1s, the splice silently changed the runtime

Why this exists

The previous workflow (audio-splice-workflow.md in krystie-profile) was a 5-step manual recipe that's already broken twice. Each fix took 30+ minutes of bash. This package makes the same workflow a one-line command.

Performance

  • scan on a 27-min audio: ~50s (Opus decode is the bottleneck on this CPU)
  • slices on a 27-min audio with 5-sec slices: ~50s
  • auto on a 27-min audio: ~10-20 min with tiny.en model on CPU. Use base.en or small.en for better accuracy at 2-4x the time. Whisper tiny.en is the right model for finding a known bad-word list — you don't need higher accuracy than "did the word 'fuck' appear at all."
  • normalize on a 27-min audio: ~30-60s (two-pass EBU R128 measurement + apply)
  • denoise on a 27-min audio: ~45-90s (FFT pass over the whole file)
  • silence-stats and threshold-tune on a 27-min audio: ~5s (just runs silencedetect with multiple thresholds)
  • batch: adds ~5s of subprocess overhead per file on top of the operation cost

Development

git clone https://git.sami/sami7777/sermon-clean.git
cd sermon-clean
pip install -e ".[test,auto]"
pytest    # 61 tests, ~17s

License

MIT — see LICENSE.

S
Description
Find bad segments in sermon audio, cut them out, splice in ElevenLabs replacements.
Readme MIT 142 KiB
Languages
Python 100%