CutNoodle for podcasters
Podcast editing is mostly subtraction. The question is which silences are dead weight and which ones are the conversation breathing, and that judgement is yours: a tool can only find the candidates and hand them to you.
Where a podcast actually loses time
Solo episodes and monologue segments are full of gaps that serve nobody: the pause while you find the next line, the sip of water, the long silence before an answer that you eventually restarted. In a video podcast these are visible as well as audible, which makes them worse.
Interviews are a different shape. Two people talking rarely produce long true silences, and the pauses that do appear are often meaningful: the beat after a hard question, the moment a guest decides how honest to be. Cutting those flattens the conversation. An automatic pass will find less to remove in an interview, and that is the correct outcome rather than a failure.
A pass that respects the conversation
- Start conservative. Set the minimum pause well above what feels aggressive. You are looking for dead air, not for every breath.
- Keep a generous margin. Speech that begins immediately after a cut sounds clipped even when the words are intact. A little air on each side is what keeps a join invisible.
- Review with sound, not just eyes. A waveform cannot show you that a pause was doing emotional work. Playing the joins is where you catch it.
- Put pauses back where they belong. Merging two clips on the timeline reinstates the gap between them in one action, which is faster than trying to find a setting that would have kept it.
Why local processing suits interview material
Guest recordings carry obligations. An interview may be under embargo until a launch date, a guest may have agreed to be recorded but not to have the raw file sitting on a third party service, and some conversations simply should not be uploaded anywhere at all. A tool that never transmits the file removes that question rather than answering it with a policy page.
The practical side matters too. Video podcast recordings are long and large, often an hour or more of multi-camera footage. Any workflow that begins with an upload begins with a wait proportional to that size. CutNoodle reads the file from disk, so the only wait is the analysis, and that reads the audio track alone. See local video editor for the details of how that works.
Fitting it into a publishing routine
The most useful place for a silence pass is early: before the creative edit, not after it. Run it on the raw recording, review the joins, export, and take the tightened file into whatever you normally use for chapters, intro, music, and the audio-only version. What arrives at your regular editor is shorter and easier to navigate, which makes the rest of the work quicker.
The export is a standard MP4 with H.264 video and AAC audio, so nothing downstream needs to change. If you publish audio separately, extracting the audio track from the tightened video keeps both versions in sync without a second silence pass.
The limits, plainly
CutNoodle listens to loudness, not to language. It will not remove filler words, it does not produce a transcript, and it cannot generate show notes or chapter markers. It handles one video track at a time, so a multi-camera edit still needs a real editor to assemble. There is no noise reduction, no levelling beyond optional loudness normalisation at export, and no music handling.
For transcript-driven podcast editing, a speech-to-text tool is genuinely the better instrument, and CutNoodle compared with Descript says so plainly. If you want the same loudness cut as a repeatable batch job across a season of episodes, CutNoodle compared with auto-editor is the comparison to read.
CutNoodle is free and runs in your browser. Open the editor and drop a video in: nothing is installed and nothing is uploaded.

