Found by a model, not a threshold
The model behind it is trained on breaths. A gate tuned to a frequency band cannot tell an inhale from a soft consonant, which is why gated narration ends up with clipped word beginnings.
Every inhale located and taken out at a strength you set — including the ones a synthetic voice inserted. It runs on your own machine, so the recording never leaves your desk.
The model behind it is trained on breaths. A gate tuned to a frequency band cannot tell an inhale from a soft consonant, which is why gated narration ends up with clipped word beginnings.
Breaths can be softened rather than deleted. Speech with every breath cut to digital silence sounds wrong to a listener long before they can say why.
Synthetic voices insert breaths that were never drawn, in places a person would not have drawn them. Nothing gives a generated track away faster.
An hour of narration holds several hundred inhales. Cutting them one by one in an editor is a morning of work, and cutting them with a noise gate takes the start of every sentence with them.
The tool that does this properly has to know what a breath sounds like. That is a model, and until recently it was a model you rented by the hour from a website you had to upload your client's audio to.
Text-to-speech systems add inhales to make narration sound human. They often add them mid-clause, at a volume no person breathes at, and identically every time — which is precisely what makes a generated voice recognisable.
Breath removal is the single pass that fixes it. Run it over the exported narration before it goes anywhere near a timeline, and the result stops announcing where it came from.
The processing happens on your own processor. No upload, no queue position, no per-minute meter, and no copy of a client's recording sitting on somebody else's server while you wait.
For anybody working under an agreement that forbids third-party services, that is not a convenience. It is the difference between being able to use the tool and not.
That failure belongs to noise gates, which act on volume. This finds breaths with a model trained on breaths, and you can soften rather than delete them.
Yes. Processing time is charged by the length of the media, not by how many tools you ran, so an hour-long episode costs an hour whether you removed breaths alone or ran the whole chain.
No. Every file is opened, processed and written back on your own machine. The app contacts our server only to check a licence.
Every tool below runs in the same job, on the same file, in the order that actually sounds right.
AI noise removal that runs offline on Windows
Remove silence and dead air from audio automatically
LUFS loudness normalisation for podcasts and video
Try it on your own recording
The free tier runs the same models on the same machine, with no account and no licence key. Install it and put a real file through it.