On this page
For a natural RX 12 Dialogue Isolate result, work on a copy and turn Instant Process off before selecting a representative difficult line in the Editor. Leave Voice near its original level, lower Noise and Reverb separately, then adjust Sensitivity if unwanted sound still clings to speech. Stop when consonants soften, ambience pumps or the voice turns phasey. Advanced adds Best / Offline quality and multiband controls; compare them on the same untouched audio before committing.
The useful target is intelligible dialogue that still belongs in the scene. A little stable room tone often cuts more naturally than maximum separation. The RX 12 controls in this tutorial were checked on September 5, 2026.
What Dialogue Isolate separates
The current RX 12 Dialogue Isolate manual says the module uses a machine-learning model trained on speech, noise and reverberant material. It divides the input into three components: Voice, Noise and Reverb. Their gain controls rebalance those detected components in decibels.
It is designed for spoken dialogue against constant or changing backgrounds: hiss, crowds, traffic, footsteps and weather are official examples. Reverb separation targets the reflections of spaces such as tiled rooms, halls and other reverberant interiors. That combination makes Dialogue Isolate useful when the background changes too much for one static noise profile.
Separation is an estimate. Do not expect it to recover words missing from the recording or reliably reveal speech completely buried under another sound. There is one Voice category, not a selector for a named speaker: overlapping people may remain together. If a critical word cannot be understood, compare another microphone or take, or request a pickup instead of guessing it from a processed artifact.
Standard versus Advanced controls
| Control | RX 12 Standard | RX 12 Advanced | Purpose |
|---|---|---|---|
| Voice, Noise, Reverb gain | Yes | Yes | Balance the three detected components |
| Sensitivity | Yes | Yes | Change what counts as noise/reverb versus voice |
| Good / Real-time | Core mode | Yes | Faster processing; still adds plug-in latency |
| Best / Offline | No | Yes | Higher-quality offline processing |
| Four-band processing | No | Yes | Scale reduction by frequency region |
Dialogue Isolate is available in RX 12 Standard and Advanced as an Editor module and plug-in; Elements does not include it. The current edition table and manual distinguish Standard's core separation from Advanced's Quality menu and multiband panel. Standard can still render a file offline in the Editor using its core algorithm; the Advanced-only feature is the higher-quality Best algorithm. The RX edition comparison covers the other tools that may justify an upgrade.
Voice, Noise and Reverb are gain controls
Voice changes the level of audio identified as speech. Keep it around the original level while judging cleanup. Raising Voice can make a processed example seem clearer simply because it is louder.
Noise changes the level of detected background noise. Lower it gradually. Maximum attenuation can expose musical noise, lisped consonants and unstable ambience that were less distracting than the source.
Reverb changes the level of detected room reflections. Treat it separately from Noise. A dry voice may be useful for a new mix, but removing all room character can make adjacent production dialogue sound mismatched.
Use matched-level A/B listening. Bypass the process, match perceived voice level, then compare words, breaths and room transitions—not just the loudest vowel.
Sensitivity changes classification, not output level
Sensitivity decides how broadly RX classifies the input as Noise and Reverb. Increasing it identifies more audio as those unwanted components and less as Voice. The manual warns that this can produce stronger reduction while adding artifacts and reducing dialogue clarity.
Set component gains first, then move Sensitivity. If traffic or room tail remains attached to speech, raise it a little. If fricatives disappear, breaths flutter or words lose body, lower it. Do not compensate for a classification error by boosting Voice; that can make damaged speech louder without restoring detail.
No universal number fits a close lav in steady HVAC, an off-axis camera mic in a kitchen and a street interview. In the Editor, add two or three settings to Compare Settings on the same difficult line. Name them for the change, such as less reverb or lower sensitivity, so you can identify the useful adjustment.
Nine steps for a natural Dialogue Isolate pass
- Preserve the source. Keep an untouched audio-file copy outside the working edit. A duplicate DAW playlist can still reference the same media. Open the working copy in the RX Editor and listen to the original.
- Disable Instant Process, then select. Turn Instant Process off before drawing a selection. Choose a difficult line with speech, changing noise, room decay and the transition into room tone; include the intended channels and full frequency range.
- Open the Editor module. Open Dialogue Isolate on that working file. This sequence uses the Editor; plug-in comparison and offline rendering use the host's controls.
- Set Voice conservatively. Keep Voice at 0 dB initially, then match perceived speech loudness when comparing the original and processed versions.
- Lower Noise gradually. Reduce distraction and compare consonants, breaths and gaps. Restore some Noise if ambience flutters or speech becomes watery.
- Lower Reverb separately. Reduce the room tail while retaining enough room character for adjacent dialogue. Back off if word endings become hollow.
- Tune Sensitivity. Raise it cautiously if noise or reverb remains; lower it if speech clarity suffers. Recheck the component gains after changing classification.
- Compare on unchanged audio. Use Compare Settings and Preview to audition alternatives. In Advanced, compare Good / Real-time with Best / Offline and adjust bands only where useful. Undo any test renders before another comparison.
- Apply once and export a new file. Return to the untouched working audio, select the full intended range and channels, and Render the chosen settings once. Export under a new filename and listen to the entire output in context, checking edit boundaries, sync and clipping.
A 10–20 second excerpt is a practical starting point, not an RX requirement. Include the quietest word and loudest interruption you need the settings to handle; later compare a contrasting passage too.
The Instant Process documentation explains why selection must come after disabling it: an enabled Instant Process immediately applies its selected tool when you make a new selection.
The Editor's Compare Settings workflow lets you audition alternatives without repeatedly processing and undoing the file. Its Render button does apply processing. Footer controls are absent in Module Chain and in RX plug-ins, with the plug-in Bypass exception. Use the DAW's bypass or snapshots for plug-in A/B; do not search for the Editor's Compare window inside the plug-in.
Good/Real-time versus Best/Offline
Good / Real-time is optimized for speed and lower latency. That makes it useful for editing and DAW playback, but the name does not promise delay-free live monitoring. In the RX 12 latency discussion on the NI forum, Jeremy_NI explains that Dialogue Isolate adds latency, requires DAW playback compensation for sync and is unsuitable when live work needs the absolute minimum delay. The latency reported by one user's host is not a universal figure for every setup.
Best / Offline is an Advanced-only quality option for the Audio Editor and offline DAW bounces or renders. iZotope describes it as offering better isolation and fewer artifacts at the cost of longer processing and higher latency. Compare it on your recording; the quality label cannot decide whether the room sound and performance still suit the edit.
Use Compare Settings on the same original excerpt for both modes. Match perceived voice level and judge sibilants, plosives, breaths, word endings and room decay. If Good preserves the performance better, use it. If you tested with Render instead, undo every test or reopen the untouched working copy before processing the full range once. Rendering a whole file over an already processed test section would clean that section twice.
Use Advanced multiband controls sparingly
Advanced divides processing into Low, Low Mid, High Mid and High bands. Each band scales the Noise and Reverb reduction from 0 to 100 percent, with three adjustable crossover points. At 0 percent, that frequency region receives none of the requested component attenuation; at 100 percent, it receives the full amount.
This is useful when a global setting fixes traffic rumble but chews up upper consonants. Keep stronger low-band processing and reduce the High Mid or High percentage until clarity returns. Conversely, a bright hiss may need more high-band attention while low voice weight stays protected.
Difference metering shows where gain is changing, but it is not a quality score. The expanded panel is highlighted when a band differs from 100 percent. Reset forgotten band values before blaming the main sliders for an unexpected result.
Dialogue Isolate versus Voice De-noise and Spectral De-noise
Dialogue Isolate separates speech from noise and reverb. Voice De-noise instead uses adaptive or learned noise thresholds and is designed for efficient, zero added processor latency. That does not remove interface or DAW buffer delay. It is a useful first comparison for modest voice noise. Spectral De-noise offers a learned or adaptive noise profile and finer tonal and broadband control, including on non-dialogue material.
For stationary hiss, buzz or line noise, the RX 12 manual itself suggests trying Spectral De-noise if Dialogue Isolate does not produce an acceptable result. The Voice De-noise versus Spectral De-noise guide explains that profile-based branch.
Try Dialogue Isolate for traffic, crowds or mixed noise and room echo. Compare a targeted tool when the problem is clearly hum, clicks, plosives or rustle. There is no compulsory chain in which every denoiser and de-reverb module must run: compare alternatives on copies before stacking them, then add only the repair still needed.
Why the cleaned voice sounds watery or synthetic
Noise attenuation is too deep. Restore some Noise gain. A quieter stable bed often masks separation residue better than absolute silence.
Sensitivity is too high. RX is classifying speech edges as noise or reverb. Lower Sensitivity and accept more background.
Reverb reduction is stripping speech decay. Restore some Reverb or use less reduction in the band where the voice becomes hollow.
The source itself limits separation. In Nick Lear's RX 12 shootout, a heavy-traffic sample had fewer vocal artifacts than RX 11 but still produced occasional booming reverb residue. His reverberant phone sample did not show the same improvement. These are observations from his recordings, not a promise about yours; compare another microphone or take when cleanup cannot preserve the words.
Too many processes are stacked. Bypass later EQ, compression and limiting while diagnosing. Restore intelligibility before polishing tone.
Keep room tone and edit transitions believable
Dialogue usually lives in a sequence, not a solo button. If one line becomes perfectly dry while the next carries the set, the cleanup calls attention to the edit. Process adjacent lines consistently or add matched room tone after repair.
Listen from at least a second before the line through a second after it. Pumping often hides inside speech and becomes obvious when the speaker stops. Preserve handles in a DAW round trip; the RX Connect Pro Tools workflow covers the selection, return and Render steps.
For isolated intrusions—a door slam between words or one bird chirp—broad Dialogue Isolate may be unnecessary. Select the event and use Spectral Repair. The spectrogram tutorial helps distinguish a narrow event from a full-band noise bed.
Stem Split is for separate components, not the final mix
RX 12 can split the selected audio into Voice, Noise and Reverb in a dedicated Stems tab. Use this when you need to inspect or process a component separately. Stem Split ignores the module's gain sliders and creates the components at 0 dB, so lowering Noise and Reverb beforehand will not bake those reductions into the stems.
For the balanced dialogue result you already approved, use the module's Render workflow instead of splitting and rebuilding it. If you do need stems, the documented export options distinguish three outputs:
- Select Voice in the tab's stem selector, then use File > Export to export that stem alone. Export Selection exports only its selected audio.
- With All Stems displayed, Export writes separate files for the stems.
- Export as Single File exports the sum of all stems. Do not treat an auditioned or active Voice lane as proof that this command exports only Voice.
All Stems view also disables module Compare and Stem Split controls. Switch to an individual stem if you need those functions. Check the actual exported file: one Voice category may still include multiple speakers, and separating a short excerpt will not create a full-length interview.
A final quality gate
- Every word remains intelligible, including consonants and phrase endings.
- Noise is less distracting without audible gating or flutter.
- Reverb is controlled without making the voice hollow or detached.
- Voice level is matched for honest bypass comparison.
- Room tone remains continuous across edits.
- No band is accidentally left at a stale multiband value.
- The entire exported clip has been heard, including passages beyond the test line.
- The rendered file retains sync, channels and headroom.
Use File > Export to write a new WAV or AIFF deliverable with the required settings. Ordinary Save overwrites an open uncompressed WAV or AIFF source. Save RX Document preserves editable history for RX, but an .rxdoc is not the audio file to send to a DAW or client. Reopen the exported audio and confirm duration, intended endpoints, channels, sample rate, headroom and sync against the original.
The broader RX background-noise guide helps when the correct module is still unclear, and the Clean Dialogue hub links the next repair by symptom.
Run one final check on ordinary speakers as well as headphones. Headphones reveal flutter and high-frequency damage; small speakers reveal whether the words remain clear without the low ambience that made the studio monitor test feel natural. If the line only works when soloed, return it to the scene with music and effects before approval. A modest repair that survives the actual mix is more useful than an impressive isolated demo that exposes a new artifact every time the background drops.
Frequently asked questions
Is Dialogue Isolate in RX 12 Standard?
Yes. Standard and Advanced include the module and plug-in. Advanced adds Best/Offline quality and the four-band processing interface.
What should I set Voice to?
Keep Voice near its original level while judging cleanup. Change it only for an intentional gain move, and match loudness when comparing bypassed and processed audio.
What does Sensitivity do?
It changes what RX classifies as Noise and Reverb instead of Voice. Higher values can remove more background but may add artifacts and reduce dialogue clarity.
Should I mute Noise and Reverb completely?
Usually not as a starting point. Lower them gradually and stop when the dialogue works in context. Some stable ambience can sound more natural than a fully isolated voice.
Is Best/Offline always better?
It is Advanced's highest-quality mode and is designed for better isolation with fewer artifacts, but every source still needs comparison. Good/Real-time may preserve a difficult voice more naturally.
Why does Dialogue Isolate sound watery?
Noise or Reverb reduction may be too deep, Sensitivity may be too high, or the source may be severely masked. Restore ambience, lower Sensitivity and compare a less aggressive pass.
When should I use Spectral De-noise instead?
Try it for stationary hiss, buzz or line noise, especially when Dialogue Isolate damages speech. Spectral De-noise also supports non-dialogue sources.
Does Stem Split keep my Dialogue Isolate slider levels?
No. RX 12 creates Voice, Noise and Reverb stems at 0 dB and ignores the module gain positions. Use Render for the approved balance. To export only Voice, select that stem and use Export; Export as Single File sums all stems.



