On this page
For a natural first pass in iZotope RX 12 Breath Control, choose Offline, Target and Natural; select a representative phrase of at least three seconds; then raise Sensitivity only until the distracting breaths respond. Do not begin at -Inf, and do not judge the module from a breath in isolation. The goal is to stop the intake from pulling attention away from the line, not to make the singer or narrator sound incapable of breathing.
RX 12 rebuilt Breath Control with machine learning, so screenshots and recipes from RX 6, 8 or 11 can send you toward controls that no longer exist in the current layout. The present module adds Real-time, Offline and Classic algorithms plus Natural/Gated ambience behavior. This guide follows the current RX 12 manual and treats the old workflow only as historical context.
Quick settings by use case
| Source | Starting route | What to protect |
|---|---|---|
| Lead vocal | Offline + Target + Natural | Phrasing, emotion and soft consonants |
| Podcast conversation | Offline + Target + Natural | Room tone and conversational rhythm |
| Tight commercial voiceover | Offline + Target; test Gated carefully | Word starts and consistent silence |
| Live DAW playback | Real-time + Target + Natural | CPU headroom and automation |
| One severe gasp | Local Gain edit or manual clip gain | The neighboring syllables |
| Old session recall | Classic only if needed for continuity | Comparing against the approved legacy render |
These are routing choices, not numeric presets. Sensitivity and Level have to follow the performer, microphone, room and processing chain. A whispery singer, close voiceover and distant interview do not present the same harmonic evidence to the detector.
What changed in RX 12 Breath Control
The current RX 12 Breath Control documentation labels the tool “STD & ADV | Module & Plug-in.” It uses machine-learning algorithms to identify breaths in dialogue or vocals by their harmonic structure rather than switching only when audio crosses a level threshold.
iZotope's RX 12 feature page calls the module entirely rebuilt and says the new neural nets improve accuracy. The practical update is the algorithm menu: Real-time for the fast plug-in path, Offline for the best documented result with more resources, and Classic for the RX 11 legacy behavior. Natural and Gated modes govern what happens to surrounding ambience.
This version is included in RX Standard and Advanced, not Elements. If a tutorial shows only Gain, Target, Sensitivity and “Output Breaths Only,” it is showing the older interface. Do not spend an hour hunting for that switch in the rebuilt module.
How Breath Control detects a breath
A gate watches level. Breath Control analyzes the selected audio and looks for breath-like harmonic structure, so it can react to a loud inhale and a quieter one without relying on a single threshold. That different detection method is useful on varied speech, but it does not guarantee correct classification. Listen for airy consonants, whispers or phrase starts being reduced along with the breaths.
The detector needs context. iZotope says Real-time and Offline tend to perform better on clips or selections at least three seconds long. A selection limited to one brief inhale may give the detector too little context. Three seconds is a recommendation for more reliable analysis, not a requirement to reduce every sound within that selection.
Use a short phrase containing several event types: a quiet breath, a loud breath, an “h” or “s” sound and the start of a word. Use that selection to tune Sensitivity, then check the rest of the recording before delivery; a good result on one phrase does not validate an entire take.
Target vs Gain: choose the right reduction logic
Target mode sets the desired destination level for each detected breath. Loud breaths are reduced more; quiet breaths that already sit below the target can remain. It is a useful first choice for songs, podcasts and audiobooks when loud breaths need more reduction than quiet ones. Target can bring several loud breaths toward the same destination; Gain instead preserves their relative level differences while applying a fixed cut.
Gain mode applies the same absolute reduction to every detected breath regardless of its starting level. It is useful when all breaths are consistently too forward, or when one selected gasp needs a predictable cut, provided the detector identifies it correctly. Manual clip gain is the option that does not depend on breath detection. The official control reference warns that heavy Gain settings can make quiet breaths vanish while loud breaths merely fall toward normal.
The documented Level range runs from 0 dB, no reduction, to -Inf, full removal. Read the mode before entering a value: in Gain it is the amount to subtract; in Target it is the destination level for detected breaths. For example, a -6 dB Gain setting reduces detected breaths by 6 dB, whereas a -6 dB Target leaves a breath already below that destination alone. This illustrates the control logic, not a recommended preset. Move toward stronger reduction while listening to the passage. If you cannot hear a problem in the mix, a more impressive soloed reduction is not an improvement.
Natural vs Gated: decide what happens to the room
Natural mode preserves ambience signals. It usually gives dialogue and vocals the smoother first pass because the room does not collapse every time the performer inhales. The remaining breath may be quieter, but the air around it stays connected to the phrase.
Gated mode removes ambience signals around the detected events. It can suit a deliberately dry commercial read or a production where all gaps will later receive controlled room tone. On location dialogue, Gated can create obvious holes that are more distracting than the breath.
The Natural/Gated switch is unavailable with Classic because it belongs to the rebuilt algorithms. If Gated is necessary, listen across cuts and under compression. You may need to fill the pause with matched ambience; the Dialogue Isolate guide explains the neighboring noise-and-room decision rather than turning Breath Control into a room-tone generator.
Real-time, Offline or Classic algorithm?
Real-time is the fastest mode and is intended for live plug-in use. Choose it when the editor or mixer must keep settings adjustable during playback. It is also the practical route for automation across a long session, but the final choice should still be checked after a full render.
Offline is the best-quality mode documented by iZotope and uses more resources. Use it in the standalone Audio Editor or an offline host when you are committing a vocal, narration chapter or repaired clip. Give the detector a long enough selection and compare the result before rendering the rest.
Classic is the RX 11 legacy algorithm. Try it against an older approved render, or when the new model behaves poorly on an unusual source. The legacy algorithm alone does not guarantee an identical session recall; retain the original settings and reference audio. It is not automatically safer just because it is familiar; it also lacks Natural/Gated control.
Step-by-step natural breath reduction
- Duplicate the file, playlist or take and keep the original untreated.
- Select at least three seconds containing quiet and loud breaths plus soft consonants.
- Choose Offline for a committed edit or Real-time for an adjustable plug-in.
- Start in Target mode so already-quiet breaths can remain.
- Choose Natural before testing Gated.
- Begin with light Level reduction and moderate Sensitivity.
- In the Audio Editor, Preview the complete phrase; in the plug-in, play the phrase from the DAW.
- Raise Sensitivity until intended breaths respond, then back off if word edges change.
- Use the Audio Editor’s Compare window to audition restrained candidates. In a DAW, compare separate protected renders or use plug-in bypass, keeping listening levels comparable.
- Render the least aggressive useful version, fix outliers manually, then export a separate delivery file and listen through it.
The Audio Editor’s Preview, Listen and Compare controls help evaluate a repair before committing it. That footer is absent from RX plug-ins except for Bypass, and is unavailable inside Module Chain. A broad pass can handle repetitive breaths, but a detection miss is a reason for a local correction, not automatically for harsher settings across the recording.
How to set Sensitivity without cutting words
Sensitivity controls how willing the detector is to classify material as breath. Too low leaves obvious inhales. Too high can grab airy consonants, whispered syllables or the noisy attack of a word. The right value is the lowest one that consistently catches the distracting events in your representative passage.
If word starts soften, first lower Sensitivity so the detector stops treating them as breaths. Moving Level toward less reduction can also limit damage, but it does not fix the misclassification. Check the unprocessed take before attributing a weak consonant to this module. If the detector still misses one breath after the rest sound correct, select that event with its surrounding context and use a local render or manual clip gain.
Old videos recommend “Output Breaths Only” to monitor detection. That control appears in the RX 6 documentation, but the current RX 12 control list does not include it. The current manual illustration instead shows a Listen button with an ear icon in the Audio Editor footer. Enable Listen, then Preview to hear the signal being removed; if you hear wanted consonants or syllables, reduce Sensitivity or repair locally. In Gated mode the removed signal can include ambience, so it is not necessarily a pure breath track. Switch Listen off before judging the normal processed output, then use Bypass and Compare on complete phrases. The plug-in does not have this editor footer: use DAW playback and bypass there.
Vocals: preserve phrasing before polish
A sung inhale often tells the listener how a phrase is shaped. Remove all of them and the vocal can feel assembled rather than performed. Use Target and Natural to lower only the breaths that jump forward after compression. Leave the quiet intake before an emotional line unless it competes with the music.
Run breath control before strong compression if the compressor is lifting the inhales, but judge again after the vocal chain. If Mouth De-click or De-plosive processing is also needed, diagnose each defect separately. A breath is broadband air; a mouth click is a short saliva impulse; a plosive is a low-frequency pressure burst. The RX vocal-cleanup workflow shows the complete order.
For a breath that overlaps a sung syllable, automated detection may not offer a clean boundary. Clip gain turns down everything in its time range, including the wanted note. A small spectral gain edit may help only when the breath occupies a separable region; neither method can guarantee clean separation of overlapping sound. Leaving a little breath is preferable to damaging the syllable.
Podcasts, voiceover and audiobooks
Long-form speech needs consistency more than surgical silence. A podcast can tolerate natural breathing; it struggles with a close-mic gasp that repeats every sentence. Audiobooks need pacing and human presence, but a narrator's loud intake can become fatiguing across hours. Commercial voiceover may call for tighter gaps, especially when music and timing are dense.
Process a representative minute before an entire episode or chapter. Check quiet speech, excited speech and edits between takes. If the room changes, one preset may create different gaps. The current iZotope vocal-cleanup guide also treats breath work as one stage among clicks, plosives and broad noise rather than the whole cleanup.
For repeated files, save a starting preset only after the microphone, distance and room are stable. Then listen through each processed output, with extra attention to quiet words and changed rooms. The RX batch-processing guide covers separate input and output locations; its initial sample check is preparation for a batch, not final acceptance of every recording.
Plug-in or standalone Audio Editor?
Use the plug-in when the session must remain adjustable, the source is long, or automation will follow changing performance. Real-time algorithm is designed for that route. Keep CPU headroom and render a short proof before committing a long delivery.
Use the standalone editor when visual context and local exceptions matter. You can select a section, run Offline, compare versions in History and apply manual gain only where detection fails. iZotope’s 2019 plug-in versus standalone guidance suggests trying automation for stray offenders before moving to manual edits. That is historical workflow advice; use the current RX 12 manual for controls and host support.
The general Editor, plug-in and RX Connect guide maps the round trip without sacrificing the original take.
When Breath Control sounds unnatural
| Symptom | Likely cause | First fix |
|---|---|---|
| Starts of words disappear | Possible false detections; check the untreated take | Lower Sensitivity and repair missed breaths locally |
| Quiet breaths vanish, loud ones remain | Aggressive Gain mode | Try Target mode |
| Room tone pumps between phrases | Gated mode or excessive reduction | Use Natural and reduce less |
| Detection misses a single gasp | Source differs or selection is too short | Include context, then use local Gain or clip gain |
| Distracting breaths remain too prominent | Sensitivity or Level is too conservative | Change one control at a time in Preview |
| New mode is worse than an old job | Model/source mismatch | Compare Classic for continuity |
Do not solve false positives by deleting more audio. Restore the original, reduce the detector's reach and treat the few remaining breaths by hand. The finished performance should pass in context with music, picture or room tone—not merely look tidy in the waveform.
Breath Control vs manual clip gain
Breath Control is the speed tool for repeated ordinary events. Manual clip gain is the precision tool for exceptions. On a three-minute vocal with consistent breaths, automation can save many edits. On a 15-second line with two complicated inhales touching words, drawing two gain moves may be faster and safer.
When automatic reduction handles most events cleanly, use a hybrid pass: let Target/Natural reduce the obvious set, then manually restore false positives and shape the outliers. Use short fades around manual changes, keep room tone continuous and compare the finished phrase against the untreated performance.
Do not replace a breath with digital zero unless the production wants hard silence. If removal exposes a hole, paste suitable room tone from the same scene with short crossfades. Breath Control does not generate replacement ambience. Strong breath reduction is an editing decision, not proof that the detector failed.
Processing order with clicks, plosives and noise
Clean events that confuse detection before committing breath reduction, but avoid an enormous destructive chain. Start with Mouth De-click for short saliva noises. De-plosive owns low-frequency P/B bursts. Breath Control owns inhalation and exhalation level.
After event repair, tackle broad hum, room noise or reverb with the module that matches it. Strong denoising can alter the texture of breaths, while strong compression can make them more prominent, so audition the full chain and revise rather than trusting a fixed order. The RX Module Chain guide provides a repeatable compare-and-stop method.
Final quality check
- Compare against the untouched take at matched loudness.
- Listen for missing H, S and F sounds at word starts.
- Check that quiet breaths still support natural phrasing.
- Check room tone through every reduced event.
- Audition before and after compression or limiting.
- Spot-check edit boundaries and crossfades.
- Test a short export, then listen through the final delivery outside RX.
- Keep the source and the approved settings for revision.
The broader Clean Dialogue hub connects breath work with clicks, plosives, hum, noise and reverb without treating every pause as a defect.
Frequently asked questions
What are the best RX 12 Breath Control settings?
Start with Offline, Target and Natural on at least three seconds of representative audio. Use light Level reduction and the lowest Sensitivity that catches distracting breaths; there is no universal numeric preset.
Should I use Target or Gain mode?
Target is usually more natural because it pushes breaths toward a destination level and can leave quiet ones alone. Gain applies the same cut to every detected breath and suits consistent or selected severe events.
What is the difference between Natural and Gated?
Natural preserves surrounding ambience. Gated removes ambience around detected breath events and can create holes unless the production wants tight silence or you replace the room tone.
Should I choose Real-time or Offline?
Use Real-time for adjustable plug-in playback and automation. Use Offline for the best documented algorithm when committing file repair; it requires more resources.
Why is Breath Control cutting words?
If the untreated word is intact, the detector may be classifying an airy consonant or phrase start as breath. Lower Sensitivity, audition a longer selection and repair remaining inhales locally. Less Level reduction can limit damage but does not correct a false detection.
Where is Output Breaths Only in RX 12?
That old checkbox is absent from the current control list. In the RX 12 Audio Editor, use the ear-shaped Listen button with Preview to hear removed signal; switch Listen off to judge the normal output. The editor footer is absent from the plug-in, where you should use DAW playback and bypass.
Should breaths be removed completely?
Usually not. Reduce breaths that distract, but preserve those that carry phrasing or realism. Complete removal can make vocals mechanical and expose discontinuous room tone.
Is Breath Control included in RX Elements?
No. The current RX 12 comparison lists Breath Control in Standard and Advanced, as both a module and plug-in.



