On this page
Background-noise cleanup in RX 12 starts with a listening decision, not a preset. If the noise is a stable hiss or fan and you have a clean section to sample, start with Spectral De-noise. If the recording is speech and the floor drifts as the speaker or microphone moves, try Voice De-noise. If traffic, crowd sound, weather or room reflections keep changing, Dialogue Isolate is a useful first test. Use De-hum, De-wind or Spectral Repair when the problem is more specific than “noise.”
The aim is rarely a silent waveform. The useful target is clearer dialogue with consonants, breaths and room perspective still intact. A quiet gap followed by a watery voice sounds more processed than a modest, consistent noise floor.
Check your edition: the editor steps below require RX 12 Standard or Advanced. Elements users can follow the Voice De-noise route as a plug-in in a DAW. Spectral De-noise, Spectral Repair and Dialogue Isolate are in Standard and Advanced; De-wind and Ambience Match require Advanced. The current RX 12 edition matrix is the reference for availability.
Identify the noise before opening a module
Listen once without touching a control. Write down whether the unwanted sound is steady, moving, pitched, intermittent or tied to the speaker. Then look at the spectrogram. Broadband hiss appears as a diffuse bed; electrical hum creates horizontal lines at a fundamental frequency and its harmonics; a short beep or clatter occupies a limited time region; reverb trails follow the words rather than existing as a separate floor.
That distinction matters because the RX 12 Spectral De-noise manual covers tonal and broadband noise, while Voice De-noise is designed for noisy vocal recordings in real time. iZotope’s own background-noise workflow separates steady and intermittent problems and sends isolated events to a spot-repair tool instead of applying one processor across the whole recording.
| What you hear | Best first test | Why |
|---|---|---|
| Air conditioner, preamp hiss or stable room tone | Spectral De-noise | It can learn a representative noise-only profile and gives detailed tonal/broadband control. |
| Speech over a floor that slowly changes | Voice De-noise | It is speech-focused and practical as a real-time plug-in. |
| Crowd, traffic or a complex moving background | Dialogue Isolate | It separates voice from changing noise rather than relying only on one static print. |
| 50/60 Hz mains tone and harmonic buzz | De-hum | The problem is pitched interference, not a general broadband floor. |
| Intermittent light or moderate wind bursts hitting the microphone | De-wind (Advanced) | It targets foreground wind rumble; constant background wind needs a different test. |
| One horn, beep, cough or chair scrape | Spectral Repair | A local selection avoids processing every clean second around the event. |
Protect the original and choose a test section
Work on a copy or duplicate the playlist/region in the DAW. If you are using the standalone editor, save a new file before the first render. Noise reduction becomes hard to judge after several minutes of repeated listening, so the untouched version is part of the workflow, not just a backup.
Choose a short section that contains three things: a quiet phrase, a normal phrase and a pause with the room audible. A loud sentence can survive settings that destroy the end of a whispered word. The quiet phrase is therefore the stricter test. Keep the playback level fixed and compare processed and unprocessed versions at roughly matched loudness; a level difference can bias the comparison.
If you are new to the editor, the RX 12 beginner cleanup walkthrough explains selections, Preview and Compare. If the waveform has flat-topped peaks as well as noise, deal with that separately using the De-clip diagnostic guide. De-noising cannot reconstruct clipped transients.
Use Spectral De-noise for a stable, learnable floor
Spectral De-noise is the deliberate choice when the unwanted sound stays reasonably consistent and the file contains a section where only that sound is present. The noise-only portion does not need to be a dramatic stretch of silence; it needs to be representative. Avoid breaths, word tails, music decay and handling noises. If the learned profile contains wanted audio, the module is being taught to remove the wrong thing.
- Select a few seconds of noise only. Use the same room, microphone and gain state as the dialogue you need to repair.
- Turn Adaptive mode off, then use Learn. In the editor, click Learn on the selection. In the plug-in, engage Learn while playing the noise-only passage, then disengage it. Do not keep learning when you switch to speech.
- Select the real test passage. Include quiet speech and a pause, not just the cleanest sentence.
- Begin with modest reduction. A 3–6 dB first audition is a conservative editorial starting point, not a universal recipe.
- Adjust detection before chasing silence. Use the editor module’s Listen control when available to audition removed audio. If you hear consonants or whole words there, reduce Threshold or Reduction and compare again.
- Compare quality modes on the hardest sentence. Turn Listen off before comparing the cleaned voice with the original. More processing time is justified only when it produces an audible improvement.
- Apply one restrained pass. In the editor, use Render on the test selection. With a real-time plug-in, leave the processing active for the comparison and follow your DAW’s bounce or export workflow when ready. Re-evaluate the result before deciding whether a second light pass is necessary.
- Apply the tested settings to the matching parts of the recording. In the editor, select the remaining unprocessed sections that share the same noise floor and render them. Do not process the test section twice. If the microphone, room or background changes, audition that section separately and learn a new profile or test Adaptive mode.
The RX 12 module controls guide identifies Listen by its ear icon: engage it, then use Preview to hear what the module would remove. Listen is available only in certain editor modules. RX plug-ins do not provide the editor’s Listen, Preview and Compare footer controls; their Bypass control remains available. Use bypass and matched-level comparisons when your plug-in has no removed-signal monitor.
Spectral De-noise also has an Adaptive mode for changing noise, but it uses more memory and processing than Voice De-noise. A noise-only sample is required for the learned procedure above, not for every Spectral De-noise workflow.
iZotope’s official procedure likewise recommends learning a noise-only section, previewing speech and noise together, and adjusting Threshold and Reduction to taste. It also warns that one heavy pass is less transparent than multiple lighter passes. That is useful guidance, but do not turn it into a ritual: if the first light pass solves the distraction, stop.
Use Voice De-noise when speech is the priority
Voice De-noise is the quicker route for spoken word when the background changes gradually or a clean noise-only sample is missing. It is also available as a real-time tool, which makes it convenient for a podcast or dialogue track that still needs automation and mix revisions. The Voice De-noise manual describes its Adaptive and Manual modes. Elements includes the plug-in; Standard and Advanced also provide the editor module.
For spoken dialogue, choose Optimize for Dialogue and start with the Gentle filter as a conservative test. Music mode is intended for sustained sung vocals. Surgical can remove more noise, but may change the timbre or introduce chirping artifacts. Start in Adaptive mode when the floor changes across the take. Preview the quietest useful words and raise the reduction until the noise stops pulling attention, then step back one notch. If the background is actually stable and you can capture it cleanly, turn Adaptive off and use Learn on a noise-only passage; the Voice De-noise manual recommends at least one second. Compare that learned result instead of assuming Adaptive is better. The winning version is the one that preserves the voice, not the one with the lowest pause noise. The RX 12 review covers the wider edition decision without turning this repair guide into a buying page.
A full-band gate that only closes between phrases leaves noise underneath the words. Voice De-noise works across frequency bands, so it can attenuate parts of the noise while other bands carry speech. Aggressive settings can still blur low-level articulation. For a complete editor-versus-plug-in decision, use the RX Editor, plug-in and Connect workflow guide.
For a closer comparison of the two denoisers, including when each adaptive workflow helps, see Voice De-noise vs Spectral De-noise.
Use Dialogue Isolate for complex, moving backgrounds
Dialogue Isolate becomes useful when “background noise” is really several changing sources: a café, passing traffic, wind through trees, a reverberant room or production sound recorded far from the actor. A static noise print cannot describe all of that. iZotope’s current cleanup material uses Dialogue Isolate for changing dialogue backgrounds and Spectral Repair for events that still need local treatment.
Test Dialogue Isolate on the difficult section first. Keep Voice gain unchanged initially, lower Noise in small steps, and adjust Reverb separately only when the room tail needs treatment. If speech becomes hollow or phasey, reduce Sensitivity or ease the Noise/Reverb attenuation. The RX 12 Dialogue Isolate manual warns that higher Sensitivity can reduce clarity. Best / Offline quality and four-band controls require Advanced, even though the module itself is also in Standard. Our Dialogue Isolate tutorial covers those controls in more detail.
An iZotope community question about Dialogue Isolate and the two de-noisers asks whether separate modules outperform the combined approach. Replies describe individual workflows, not a controlled comparison. Compare the modules on the same quiet phrase and retain the least damaging result.
Treat hum, wind and isolated events separately
A persistent pitched electrical line is a reason to compare a targeted De-hum pass; it does not prove that the broadband denoiser is faulty. Use De-hum for a fundamental and its harmonics. When a strong hum dominates the recording, try treating it before broadband hiss, then compare that order with a single Spectral De-noise pass. Its separate tonal and noisy controls may already be enough.
The De-wind manual limits its intended use to intermittent light or moderate bursts hitting the microphone. It is an Advanced editor module, not a general wind plug-in. Constant background wind can be preserved as part of the noise floor; try Spectral De-noise for that case. Heavy gusts that overload and distort the microphone are outside De-wind’s intended scope. Preserve intelligible speech and accept residual movement when stronger processing makes it thin or synthetic.
For a single car horn, cough, microwave beep or chair scrape, make a time-frequency selection and test Spectral Repair. The official RX audio-cleanup guide recommends Spectral Repair for sudden events and dropouts. Processing only the damaged region is usually safer than running a stronger global denoiser.
Listen for artifacts instead of chasing a number
Listen for these warning signs. They suggest checks, not a unique diagnosis:
- Watery or chirping consonants: try less reduction and check that the learned profile contains no speech; filter and artifact-control settings can also matter.
- Pumping between words: check adaptive behavior, threshold and any gate elsewhere in the chain; a static process can also produce gated noise bursts.
- A hollow, distant voice: compare the untreated phrase to check whether processing has removed room cues or vocal detail.
- A metallic tail after every sentence: compare denoising and reverb processing separately; the sound alone does not identify one cause.
- Perfectly silent gaps with noisy words: a gate or edit may be hiding pauses without repairing speech.
- Complete syllables in the removed signal: back off immediately; wanted audio is being subtracted.
Compare on headphones and ordinary speakers. Headphones expose swirls and high-frequency damage; speakers reveal whether the remaining floor is actually distracting in context. If the voice will sit under music, check it under music before doing another pass. A repair that sounds slightly conservative in solo can be the transparent choice in the finished program.
Preserve believable room tone
Do not erase every pause. Dialogue edited to digital silence between phrases sounds gated and draws attention to every cut. Preserve a stable bed from the same room or rebuild transitions with room tone. The Ambience Match manual describes an Advanced-only tool for matching the noise floor between recordings. It adds or replaces ambience; it cannot reduce ambience already in a selection. Use it to restore continuity after cleanup, not as another denoiser.
For a single clip, short crossfades and copied room tone may be enough. For a longer scene assembled from multiple takes, match the ambience after the primary repairs so you are not matching a noise bed that will later change. This is also where restraint pays: a consistent low floor is easier to blend than a voice damaged by overprocessing.
A practical order for mixed problems
When a recording has more than one fault, use the smallest chain that solves it:
- repair clipping if clipped peaks prevent reliable listening;
- remove strong electrical hum or a specific tonal interference;
- reduce the broad noise floor with one appropriate de-noiser;
- spot-repair isolated noises that remain;
- treat reverb only when it still harms intelligibility;
- restore consistent room tone across edits;
- check the result in the real mix and against the untouched file.
This is a decision sequence, not a mandatory seven-module chain. Skip every step the recording does not need. If the repair is part of a larger dialogue job, the Clean Dialogue hub keeps the specialized module guides separate so one page does not pretend every noise problem is identical.
Frequently asked questions
How much noise reduction should I use in RX?
Use only enough to stop the noise distracting from the words. A 3–6 dB first audition is conservative, but the correct amount depends on the recording. Judge quiet consonants with bypass and, where available, the editor’s Listen control. Stop if the removed audio contains recognizable speech.
Should I use Voice De-noise or Spectral De-noise?
Use Voice De-noise as a fast speech-focused test, especially for a changing floor or real-time workflow. Use Spectral De-noise when the noise is stable, you can capture a representative noise-only section, and you need more detailed tonal and broadband control.
Can iZotope RX completely remove background noise?
Sometimes it can make the noise effectively inaudible, but complete removal is not a safe universal target. When noise overlaps speech, stronger subtraction also removes vocal detail. Clear, natural dialogue with a small residual floor is often the better result.
Do I need a noise-only sample?
A clean sample is useful for learned Spectral De-noise, but it is not required for every workflow. Voice De-noise and Spectral De-noise both have adaptive modes; Spectral’s is more demanding on the computer. Dialogue Isolate separates speech, noise and reverb without a learned static print.
Why does the voice sound metallic after noise reduction?
There is no single cause. Compare the unprocessed phrase, then test less reduction, a clean learned profile, or gentler filter settings. If a separate reverb processor is active, bypass it too. Keep the version that preserves quiet speech.
Should I remove hum before broadband noise?
Try that order when a strong 50/60 Hz tone and harmonics dominate. De-hum can target the pitched interference, but Spectral De-noise also has tonal-noise controls. Compare the results rather than treating the order as mandatory.
Is Dialogue Isolate always better for speech?
No. It is useful for complex or changing backgrounds, but a light Voice De-noise or learned Spectral De-noise pass can be more transparent on a stable recording. Compare them at matched level and keep the least damaging result. Dialogue Isolate’s Best / Offline mode and four-band controls require Advanced.



