iZotope RX World
Dialogue Cleanup

Voice De-noise vs Spectral De-noise: Which RX Tool First?

By Brandon Hayes Published Sep 3, 2026 Revised Sep 4, 2026 10 min read

Short answer

Start with Voice De-noise for efficient dialogue or vocal cleanup. Choose Spectral De-noise when a learned noise profile and detailed tonal/broadband control matter more.

Dialogue waveform beside a detailed spectral noise profile in an audio restoration studio
Editorial illustration comparing voice-first and profile-based noise reduction; not an RX interface screenshot.

Start with Voice De-noise when the wanted signal is dialogue or singing and you need a fast, efficient pass that can follow a changing noise floor. Start with Spectral De-noise when the unwanted sound has a recognizable profile—hiss, fan, camera motor, electrical buzz—or when you need separate control over tonal and broadband reduction.

Voice De-noise is a speech-aware multiband processor suited to a track insert. Spectral De-noise gives you more control over a noise profile, frequency-dependent reduction and processing artifacts. Voice De-noise adds no processing latency of its own; that does not make the audio interface, host buffer or complete monitoring path latency-free. Choose the first audition by the source and noise, then judge the result in context.

Voice De-noise vs Spectral De-noise at a glance

QuestionVoice De-noiseSpectral De-noise
Best starting pointDialogue, voiceover or sung vocalAny source with steady or slowly changing tonal/broadband noise
Noise analysisAdaptive threshold or a manually learned referenceFixed learned noise profile or adaptive profile
Real-time useDesigned for high efficiency and zero latencyPossible in lighter modes; higher-quality modes are heavier and add latency
Core controlDialogue/Music, Gentle/Surgical, threshold and reductionSeparate noisy/tonal threshold and reduction, quality, artifact control and frequency curve
Edition accessPlug-in in Elements; module and plug-in in Standard and AdvancedModule and plug-in in Standard and Advanced

A useful rule is voice first, noise profile second. If the source is speech and the problem is ordinary room noise, preview Voice De-noise. If it leaves a stable hiss or tone, compare a restrained Spectral De-noise pass. Do not stack both by habit.

What Voice De-noise is actually doing

The current RX 12 Voice De-noise manual describes 64 psychoacoustically spaced bandpass filters acting as a multiband gate. Signal above each band’s threshold passes; material below it is attenuated. RX can analyze speech and set those thresholds automatically, or you can learn and adjust them.

Adaptive mode adjusts the thresholds as the noise floor changes, so you do not need isolated room tone to begin. The six visible threshold nodes control the curve; they are not 64 separate on-screen controls. The master Threshold slider offsets those nodes together. Dialogue optimization responds to short spoken bursts, while Music is designed to preserve sustained sung notes.

If the thresholds sit above quiet wanted detail, that detail can be attenuated along with the noise. Reduction sets the maximum depth of attenuation below the threshold. Surgical removes more noise than Gentle but can alter timbre and produce chirpy or watery residue. Watch the input/output spectra while listening to consonants and vocal tails; the meters do not decide whether the result sounds natural.

What Spectral De-noise is actually doing

According to the RX 12 Spectral De-noise manual, the module learns a profile of stationary or slowly changing tonal noise and broadband hiss, then subtracts that noise when the signal drops below its threshold. Typical targets include tape hiss, HVAC, fans, camera motors, line noise, ground loops and complex harmonic buzz.

Spectral De-noise separates tonal components such as hum or interference from random components such as hiss. You can set different thresholds and reduction depths for each, edit a frequency-dependent reduction curve, choose among quality algorithms and manage artifacts. That is the main reason to use it: not because “spectral” automatically means better, but because the noise needs controls Voice De-noise does not expose.

A learned profile remains fixed during processing, which suits a continuous, consistent noise bed. Adaptive mode can track changes in outdoor ambience, traffic or waves, but it uses substantial memory and processing power. Its Learn time setting is a lookahead interval used to decide what counts as noise. For an efficient adaptive voice insert, the manual points back to Voice De-noise.

Choose Voice De-noise for dialogue and vocal speed

For a podcast, interview, voiceover, livestream or sung take with modest room noise, Voice De-noise is a sensible first audition. Its efficient insert workflow lets you hear cleanup against the mix and revise settings later. Dense competing voices or strong room reflections may call for a different tool, even when the wanted source is speech.

Use Dialogue for spoken phrases and Music for sung vocals. Begin with Gentle, reduce only the distracting noise, and bypass in context. Try Adaptive when the floor changes; with a representative noise-only section, compare a learned reference instead. These are starting choices, not evidence that either mode will always preserve more detail.

Elements owners get Voice De-noise as a plug-in, which makes it the only one of these two choices in that edition. The RX 12 edition comparison explains when the standalone editor and Spectral De-noise justify moving to Standard.

Choose Spectral De-noise for a defined noise fingerprint

Quality A has the lowest processing cost and latency. B adds time-and-frequency smoothing and can still run in real time on some systems. C uses multiresolution processing to protect transients; D adds high-frequency synthesis to help recover detail masked by noise. iZotope does not recommend C or D for real-time operation. Compare them offline on the actual recording: the “Best” end of the slider is not a promise that D will sound best on every source.

Choose Spectral De-noise when you can identify representative noise, or when separate tonal and broadband controls solve a problem the simpler voice processor leaves behind. Preamp hiss on music, a constant camera motor and complex electrical buzz are useful candidates. A noise-only sample helps with fixed learning; Adaptive remains an option when no suitable sample exists.

Compare it when Voice De-noise changes the voice before reducing the contaminant enough. Learn the longest representative noise-only selection available—ideally a few seconds—then audition words, pauses and transitions. Judge consonants and sustained notes as carefully as the silence between them.

For a visual, multi-step cleanup, use the standalone editor or a deliberate round trip rather than forcing the heaviest mode onto a live insert. The RX Editor vs plug-ins vs Connect guide maps those choices without treating one route as inherently more professional.

Adaptive mode and Learn solve different problems

Learn assumes the captured reference represents the unwanted sound throughout the target. In Voice De-noise manual mode, iZotope recommends at least one second of pure noise; longer references place threshold nodes more reliably. In Spectral De-noise, the manual recommends the longest available noise-only section, ideally a few seconds.

Adaptive continually updates the estimate. Voice De-noise adjusts its threshold conservatively; Spectral De-noise updates the noise profile and uses a Learn time lookahead. Neither is a guarantee of transparent separation when noise and wanted sound overlap.

Do not use Adaptive merely because the button exists. If a refrigerator and microphone preamp create the same bed across an entire interview, a good fixed Learn can be more predictable. If the speaker walks from a quiet room to a street, one fixed profile may fail; split the clip by environment or use an adaptive pass and check every transition.

A safe Voice De-noise starting workflow

  1. Work on a duplicate audio file in the editor, or keep an untouched source and a bypassable insert in your host. Choose Dialogue for speech or Music for singing, then start with Gentle.
  2. For changing noise, enable Adaptive mode. For a fixed reference, turn Adaptive off and choose at least one second containing only representative noise. In the editor, select that region and click Learn. In an insert, engage Learn and play only the noise reference; stop playback and disengage Learn before playing the performance.
  3. Return the editor selection to a difficult phrase containing quiet detail, full-level voice and a pause. Use Preview and Bypass; with an insert, use host playback and plug-in bypass instead.
  4. Raise Reduction only until noise stops pulling attention. If quiet consonants, breath or decay thin out, reduce the threshold or reduction and compare again.
  5. In the editor, click Render to apply the chosen settings to the selected target. Preview alone does not apply them. Keep an insert editable until you deliberately bounce or render through the host.
  6. Check the result against the original at a matched listening level, both solo and in the mix. If you rendered a short test, Undo it before processing the larger passage so that the excerpt is not denoised twice. Export an approved editor result to a new file and listen to that file.

This is a listening sequence, not a universal preset. Microphone distance, room tone, compression and the final bed all change how much residual noise is acceptable. The broader RX background-noise workflow helps identify hum, reverb, wind and intermittent events before choosing a denoiser.

A safe Spectral De-noise starting workflow

  1. Keep an untouched original. Turn Adaptive mode off for fixed learning and select the longest representative noise-only region available, ideally a few seconds. Exclude wanted breaths, reverb tails and distant speech.
  2. In the editor, click Learn on that selection. In the plug-in, engage Learn and play the reference; for AudioSuite, use Preview on the noise selection. Stop reference playback and disengage Learn before auditioning wanted audio. Return the editor or AudioSuite selection to the passage you intend to process.
  3. Preview with modest reduction. Use the chain-link controls to separate Noisy and Tonal adjustments when hiss and harmonic noise need different treatment. Raise a threshold cautiously: it can remove more noise but also suppress quiet wanted detail.
  4. If frequency shaping helps, enable Reduction curve. Raise a curve point for LESS reduction in that region; lower it for MORE. Preserve vocal brightness instead of reducing all frequencies equally.
  5. Compare A–D on a short, difficult phrase; allow offline processing for C/D. In the editor, use Compare to create alternatives, select an entry in Compare Settings and Preview it against the original. This prepares alternatives without applying them.
  6. If you use the ear-shaped Listen control to hear removed material, turn it off before the final comparison and render. Recognizable wanted detail in the removed signal is a reason to back off. Adjust Artifact control only after checking the profile, thresholds and reduction.
  7. Select the winning Compare Settings entry and click its Render button to apply it, or use the module’s Render for its current settings. Undo any rendered short test before selecting and processing the larger target. In a plug-in, use the host’s render or bounce workflow; the editor footer is not part of an insert.
  8. Compare the processed target with the original, including pauses and edit boundaries. Export the approved editor result to a new file, keep the original, and listen to the exported file before replacing audio in the mix.

iZotope’s June 2026 vocal cleanup walkthrough demonstrates learning noise and shaping the reduction curve. The RX 12 Preview, Compare and Render controls explain the separate audition and apply steps. You can keep the editor’s working history in an RX document; use audio export for a file your mix session can play.

Listen for artifacts, not just remaining noise

Check consonants and sibilants, sustained vowels or notes, and transitions into silence. Chirping or a watery metallic veil can indicate musical-noise artifacts. Quiet syllables becoming smaller can indicate a threshold that is too high. Pumping or bursts of noise can arise from the gating behavior; these listening clues are not a unique diagnosis of the cause.

In Spectral De-noise, lower Artifact control values favor spectral subtraction, which can leave musical noise. Higher values favor broader gating, which can reduce chirping but leave noise bursts after the wanted signal falls below threshold. Move the control a little, compare, and undo changes that trade one distraction for another. A little steady noise under music may be less distracting than watery residue.

The question remains practical for users: a June 2026 iZotopeAudio discussion asks whether Spectral De-noise is better for spoken word and whether separate denoise/de-reverb modules beat Dialogue Isolate. The question supports comparing these routes; it supplies no controlled test or universal winner. Choose by intelligibility and natural tone on your own recording.

When neither module should be first

If an isolated defect dominates, audition a tool aimed at that defect: De-hum for harmonic hum, De-click or Mouth De-click for impulses, and De-plosive for mic thumps. De-reverb or Dialogue Isolate may fit reflections or difficult dialogue ambience. A local spectral repair may be enough for one cough or squeak. Check your edition before choosing those tools; they are not all included in Elements.

Order depends on what masks the problem. Obvious clipping or strong hum can justify a dedicated repair before finer denoising, but a lighter noise pass may reveal defects you could not hear earlier. Compare the order on a copy instead of committing a fixed chain to the whole file. The RX 12 beginner tutorial covers Preview, Compare and History so you can hear and undo each change.

The practical decision tree

  • Dialogue or sung vocal, modest variable floor, live insert needed: Voice De-noise first.
  • Dialogue with a clean sample of stable hiss or fan noise: compare Voice De-noise with a learned Spectral pass.
  • Music, effects, archive audio or production sound with stationary noise: Spectral De-noise first.
  • Changing outdoor noise and ample offline processing: Spectral Adaptive may fit, but compare its artifacts and resource cost.
  • Hum, clicks, reverb, wind or isolated events dominate: audition the relevant dedicated repair and compare the processing order.

Browse the Clean Dialogue hub when voice is the source, or the Repair Audio hub when the artifact should decide the module.

Frequently asked questions

Is Voice De-noise better than Spectral De-noise for dialogue?

It is a sensible first audition for modest dialogue noise because it is speech-aware, efficient and adds no processing latency of its own. Spectral De-noise may work better when a stable noise fingerprint needs detailed control. Neither guarantees a transparent result.

Does Voice De-noise need a noise profile?

Not in Adaptive mode. It can adjust thresholds from incoming audio. Manual mode can learn a pure-noise reference, and iZotope recommends at least one second for that capture.

How much noise should I remove?

Remove only enough that the noise no longer distracts in the final context. Back off when consonants, breath, reverb tails or sustained notes thin out, or when chirpy and watery residue appears.

Can Spectral De-noise run in real time?

Quality A is the lightest option; B can also run in real time on some systems. iZotope does not recommend C or D for real-time operation. Spectral De-noise uses more resources and adds more processing latency than Voice De-noise; audition the chosen mode in your host.

Should I use Adaptive mode or Learn?

Use Learn when a clean noise-only reference represents a stable problem. Use Adaptive when the floor changes over time, then check transitions carefully because the detector is continually updating.

What is the difference between Gentle and Surgical in Voice De-noise?

Gentle favors transparency and removes less high-frequency sizzle. Surgical removes more noise but can alter timbre and create musical-noise artifacts sooner.

Does RX Elements include both denoisers?

No. RX 12 Elements includes Voice De-noise as a plug-in. Spectral De-noise is included as a module and plug-in in Standard and Advanced.

Should I run Voice De-noise and Spectral De-noise together?

Only when a light first pass leaves a distinct second problem and each module has a clear job. Compare the stack against a single restrained pass; two denoisers can compound artifacts.