Turn a sermon into
clips worth sharing
Drop in a video or audio file — or record one here — and Clip transcribes it with Whisper in this browser, reads the waveform, and hands you clip ranges, cut points and captions — then an export that's ready for Sunday night.
Nine things it does
so you don't have to
Tap any card for the longer version.
One sermon,
start to finish
Drop the file, let the transcript arrive, mark the moments that preach, export. That is the whole loop.
How it technically works
A sermon is a file sitting on your disk. Every stage between that file and a finished clip runs inside this browser tab, on your own processor. Here is the long version, for whoever has to sign it off.
-
The file is opened, not loaded
Dropping a sermon in hands the browser a
Filereference — a handle on bytes that stay where they are. Clip reads it in ranges throughBlobSource, so a four-gigabyte service is opened in a moment and never copied, never buffered whole, and never written back. Your original is untouched by everything that follows.File APIRanged readsRead-only -
Audio is decoded packet by packet
The audio track is pulled through WebCodecs one packet at a time and resampled into a single mono 16 kHz
Float32Array— the rate Whisper wants, and plenty for reading the room.The obvious alternative,
decodeAudioData, has to decode the whole file in one call: most of a gigabyte for a 37-minute sermon, and past a certain size some browsers hand back a buffer of the right length that is silent after the first few minutes. Nothing downstream can tell that from a quiet sermon. Streaming sidesteps it; the single-call path stays as a fallback.WebCodecsmediabunny 1.4.416 kHz mono -
The waveform is measured, not guessed
Level is taken as RMS over 512-sample frames — 32 ms each — which gives the envelope you see, the noise floor, every pause long enough to be rhetorical, and the lift in the room after a line lands.
On video, the encoded packets are walked metadata-only to index every keyframe, so the places a lossless cut can legally start are drawn on the timeline before you commit to one rather than reported as an error afterwards.
RMS · 32 ms framesKeyframe indexNo model -
Whisper runs on your hardware
Transcription happens in a module worker so the interface keeps its frame rate. The model is an ONNX build of Whisper driven by transformers.js: WebGPU when
requestAdapter()answers, otherwise ONNX Runtime's SIMD WASM build — an fp32 encoder with a q4 decoder on the GPU, q8 throughout on WASM.Audio goes through in 30-second windows with a 5-second stride, returning word-level timestamps, with a segment-level pass as a fallback for checkpoints that ship no alignment heads. Every run records what it actually managed, per model and per backend, so the estimate you are shown next time is a measurement rather than a promise.
tiny · base · smallWebGPU or WASM Worker threadWord timestamps -
Clips are scored, not generated
Each transcript segment is embedded as a 384-dimension vector by MiniLM — mean-pooled and normalised — and cosine-matched against six anchor centroids built from written examples of what a quotable, pastoral or convicting moment sounds like. That score is blended with the pause before, the response after, energy, repetition, how well the length fits, and how far the candidates are spread across the hour.
No language model writes anything, and if the scorer will not load the whole thing falls back to lexical matching rather than failing. Scripture references are a separate matter entirely: a grammar-based parser reads them out of the text, so the same sermon yields the same references every time.
all-MiniLM-L6-v2Cosine + prosody bcv parserDeterministic refs -
Editing touches nothing
Clips, cut points, corrections and caption styling are arrays and objects in memory, rendered onto a canvas over a plain
<video>element. A removed passage is a range that the playhead skips and the exporter never writes.Which is why nothing here is destructive, and why every edit is reversible right up until you export: there is no working copy being progressively damaged, only a list of decisions about a file that has not changed.
In-memory modelCanvas previewNon-destructive -
Export encodes locally, or not at all
CUT copies the encoded packets straight into a new container, snapped to the nearest keyframe — no re-encode, no generation loss, the same video your camera made.
RENDER decodes frames, draws each one into an
OffscreenCanvaswith the vertical crop and captions applied, and encodes back out to MP4 or MOV through WebCodecs. Audio is copied when it can be, re-encoded when the edit forces it, and written as PCM when the platform's AAC encoder is the sort that produces a track full of silence — which Clip catches by decoding the finished file and measuring its peak before handing it to you.Lossless remuxWebCodecs encode MP4 · MOVVerified audio
Where things are kept
All of it in this browser, on this device, under your own control. Clearing site data is a complete erase — there is no second copy anywhere to fall back on.
| Store | What | How long |
|---|---|---|
| localStorage | Settings, caption style, the correction glossary, measured transcription speeds | Until you clear it |
| IndexedDB | Per sermon: transcript, clips, cut points — plus the media file itself for the two most recent, so a refresh does not cost you the afternoon | Work kept; media evicted after two |
| Cache API | Whisper and scorer weights, exactly as downloaded | Until the browser evicts them |
| A server | nothing — there isn't one | n/a |
Model weights run 42 MB to 563 MB depending on which Whisper you pick and whether your machine has WebGPU. Downloaded once, then read from cache for ever after.
What crosses the network
Every arrow points inward. Clip fetches code and model weights and sends nothing back — there is no endpoint to send it to, no account to attach it to, and no setting that changes either.
| Request | From | When |
|---|---|---|
| The page, stylesheet, scripture parser | Wherever Clip is hosted | On load |
| transformers.js + ONNX Runtime | The copy beside the page; jsDelivr only if it isn't there | First transcription |
| Whisper and MiniLM weights | Hugging Face | Once per model |
| mediabunny | The copy beside the page; jsDelivr fallback | First cut or render |
| Your audio, video, transcript, clips, stills, filenames | never sent | never |
Put the libraries in ./vendor and the last CDN dependency disappears too. Once the page and the weights are cached, transcribing, scoring, cutting and rendering never touch the network again — no analytics, no telemetry, no beacons.
Why on-device
A sermon is a pastoral document before it is content. The name of the family being prayed for, the confession made from the platform, the illness mentioned in passing — none of that belongs in somebody else's logs because you wanted forty seconds for Instagram.
Why it is also faster
The hardware is already in the room. Uploading four gigabytes to rent a GPU, waiting in a queue and downloading the result is slower than a browser that has WebCodecs, WebGPU and your own file on local disk — and it fails on church wifi, which is where this work actually happens.
Why it is free
Nothing here bills by the minute. No inference to pay for, no storage to rent, no accounts to run — so there is no business reason to meter you, and nothing held anywhere to breach, subpoena or sell.
The editor is one HTML file, unminified and commented.
Every claim in this section is checkable in view-source.
What's said in the room stays in the room
Clip reads, transcribes, scores, cuts and renders on your own machine. There's no account to make, no upload, and no server behind it — so this isn't a setting you have to trust us on. It's the only thing the app is able to do.