On this page
For a sung vocal in RX 12, start with a copy of the finished comp, identify the distracting defects, and repair clipping, clicks and plosives before tackling broad noise. Use Music optimization in Voice De-noise and Music mode in Repair Assistant for singing. Keep useful breaths and room continuity, then judge the repaired file against the original in the mix.
This workflow is for an isolated vocal recording. If you only have a finished song and need to extract its singer, start with Music Rebalance and stem separation. Cleaning an existing vocal track and separating one from instruments are different jobs.
Check which RX tools your edition includes
The Editor workflow below requires Standard or Advanced. The current RX 12 edition comparison includes Mouth De-click, De-plosive, Breath Control, De-ess, Spectral De-noise and Spectral Repair in both. Advanced is not required for this basic vocal cleanup.
Elements provides six plug-ins: De-click, De-clip, De-hum, De-reverb, Voice De-noise and Repair Assistant. It does not include the standalone Audio Editor or the dedicated Mouth De-click, De-plosive and Breath Control modules. Use its available effects in your host, with local gain edits for isolated breaths or pops when appropriate. Do not search an Elements installation for an Editor-only spectral selection tool.
Prepare a copy and listen before processing
Finish the intended comp, save the session and export or consolidate a separate working audio file while retaining the original media. A duplicated track or playlist can still point to the same file. Preserve the vocal's start position, sample rate and channel layout; retain enough surrounding audio to check edits and transitions. Keep creative effects out of the repair copy unless they are deliberately part of the source.
Listen in the mix, then solo. Nick Messitte's June 2026 vocal-cleanup walkthrough uses that sequence to distinguish distracting problems from details that disappear in context. You can briefly audition the intended compression to reveal rising breaths or room noise, then bypass it before evaluating the raw repair.
Mark overloaded phrases, lip smacks, P/B bursts, tonal buzz, hiss, room changes, headphone spill and harsh syllables. The RX spectrogram guide helps locate them: clicks often look like short vertical impulses, tonal interference like horizontal lines, and plosives like low-frequency bursts. These shapes are clues; listen before selecting a repair. Vocal harmonics also form lines.
Nine steps for a natural RX 12 vocal cleanup
- Protect the source. Save the session and create a separate audio-file copy of the intended comp, retaining timing, channels and the original media.
- Diagnose before selecting. Listen in the mix and solo, mark the defects, and turn Instant Process off before making selections in RX Audio Editor.
- Repair overload and hard glitches. Test De-clip on genuinely clipped phrases and De-click or local Spectral Repair on isolated faults. Check restored peak headroom.
- Treat plosives and mouth clicks. Preview the affected phrases, protect consonants and vowel body, and use removed-signal monitoring only where the chosen module supplies it.
- Reduce confirmed tonal noise. Use De-hum for unwanted hum or buzz; skip it when the lines belong to the singer. Learn from isolated interference when available.
- Choose the broad repair. Test one appropriate noise or reverb tool at a time. Select Music optimization in Voice De-noise for singing and preserve sustained notes.
- Balance breaths and sibilance. Reduce distracting events while keeping phrasing, ambience and intelligibility. Compare a different order if the two processes interact badly.
- Approve each stage. Compare on the same stage input, undo test renders, disable removed-signal monitoring and render the selected repair once. Spot-fix remaining defects.
- Export and check the return. Export a new audio filename, place it at the original position in a session copy and listen through the full vocal and mix for tone, sync, peaks and transitions.
This is a starting order, not a requirement to run nine processors. Mike Metlay's 2021 order-of-operations article explains the rationale for repairing clipping, clicks and hum before broad denoising. Its own breath/de-ess example changes order after listening. Current controls come from the RX 12 manual, not that historical interface.
The Compare window auditions alternative settings before commitment; it is not a substitute for a saved stage backup. Use a named RX Document or separate audio export when you need a durable checkpoint. After a stage has been rendered, do not apply the entire chain again at export.
Repair clipping and hard glitches
De-clip estimates replacement peaks; it cannot prove what the original performance contained. Select a phrase with audible clipping and inspect its histogram. Set the threshold just below the actual clipping concentration, which may sit below 0 dBFS if the file was turned down earlier. Compare the repaired phrase and its boundaries at similar loudness.
Reconstructed peaks can rise. Leave headroom using the module's gain control and check the resulting peak level before export; a 32-bit float file does not prevent overload at playback. Follow the De-clip threshold and headroom procedure for the detailed controls. Avoid printing compression or normalization before diagnosing the source damage.
For narrow digital ticks, test De-click's Single Band algorithm. Mouth noises may need Mouth De-click or another De-click algorithm; the names do not make their capabilities mutually exclusive. For one local defect, Spectral Repair's Attenuate reduces selected material, while Replace reconstructs it from its surroundings. Keep the selection tight enough to preserve the note.
Interpolate is limited to individual clicks shorter than 4,000 samples. It is not a tool for rebuilding a missing word. If another clean take exists, a careful edit may preserve the performance better than a speculative repair.
Reduce plosives before high-pass filtering
A blast of air hitting the microphone can create a low-frequency pop and sometimes overload the recording path. Select the burst with enough context to hear the following vowel. Preview De-plosive and compare bypass; its current Editor illustration does not supply the Listen ear button found in some other modules.
The De-plosive manual recommends processing before a high-pass filter because detection looks for energy between 20 and 80 Hz. That detection range is different from Frequency Limit, which sets the upper boundary of the reduction.
Sensitivity controls what is classified as a plosive; Strength controls how much the detected material is reduced. Set Frequency Limit from the actual burst rather than a preset for the singer's range. If the vowel thins or the consonant loses definition, reduce the treatment and retest. A pop filter and an adjusted microphone angle can prevent the next take from needing the same repair.
Use Mouth De-click without erasing consonants
Mouth De-click targets lip smacks and similar mouth noises. It supports longer selections and individual events. Test a representative phrase first; extend a light pass to the take only if the clicks are widespread and healthy consonants survive.
Increasing Sensitivity detects more events and can damage plosives. Frequency skew weights detection toward lower or higher frequencies; it is not a hard frequency boundary. Click widening extends the repair around each detected event, so more is not automatically better.
In the RX 12 Editor, enable the Listen ear icon and Preview to hear removed material. Older tutorials call this Output Clicks Only. If you hear wanted syllable attacks or vocal detail, back off; disable Listen before normal playback and rendering. Plug-in and Module Chain footers differ, so use the controls actually available in that context.
De-crackle can help with dense, quiet crackling after the larger clicks are handled. It is another repair to test, not a compulsory second pass. Fix remaining isolated clicks locally instead of escalating a setting that already changes the singer.
Remove hum without following the melody
Use De-hum for audible unwanted tonal interference. Static suits simple hum with a few harmonics; Dynamic can address more complex buzz, including unrelated tones. A perfect 50/60 Hz ladder is not required for every De-hum use, and a low sung note is not evidence of electrical hum.
Learn from interference alone when available. In the Editor, select it and click Learn; in the plug-in, engage Learn and feed the relevant audio, then leave learning before judging the vocal. A mixed sample with prominent hum is possible but less reliable. Compare the lowest sung notes and sustained harmonics as well as the gaps.
Adaptive is an option within the filter types, not a third filter tab. Be especially cautious with Adaptive Dynamic on sustained singing; it can mistake wanted tonal material for interference, and the manual recommends offline use of that option in the plug-in. The current De-hum guide covers filter choice, learning and ringing. Use its removed-signal monitor where available, then return to normal output.
Choose one primary noise tool
| Problem | First comparison | What must survive |
|---|---|---|
| Light steady booth noise | Voice De-noise, optimized for Music | Sustained notes and soft phrase endings |
| Hiss or buzz needing detailed control | Spectral De-noise with a suitable noise profile | Vibrato, consonants and a stable residual floor |
| Changing noise mixed with voice and room | A careful Dialogue Isolate trial | The entire sung phrase, especially breathy notes |
| Excessive room reflections | De-reverb; compare Dialogue Isolate if useful | Vowel body, natural decay and believable space |
Voice De-noise's Music optimization is intended for sustained sung notes; Dialogue optimization reacts faster to speech. Start with Gentle filtering and modest reduction. In Manual mode, Learn from at least a second of noise without singing, breath or reverb you want to retain. Finish learning and return to the vocal selection before evaluating the result.
Spectral De-noise offers separate tonal and broadband reduction. Prefer a few seconds of representative noise alone for a fixed profile, and relearn when the recording condition changes. Digital silence is not a useful recording-noise sample. Its Artifact Control trades watery residue against a more gated result; turning it up is not a universal quality improvement. The Voice and Spectral De-noise comparison explains those choices.
Dialogue Isolate is trained around spoken dialogue. It may help singing, but do not approve it from an isolated noise-free gap. Compare the singer's lowest and highest notes, falsetto, vibrato and breathy entrances. Increase reduction only while the performance remains intact; a quieter background does not compensate for missing vocal detail.
Reduce room only when it distracts
The general De-reverb module provides a documented route for roomy vocals. Learn from several seconds containing a little room tone, the direct voice and its decay. A noise-only sample is appropriate for some denoisers, but it does not provide this complete reverb example.
Preview modest Reduction and check phrase endings. If tails remain, inspect Tail Length; if the voice sounds dull or processed, back off the treatment. Reverb removal and denoising can interact, so compare their order on the same input when necessary. Do not apply both simply because a chain preset includes them.
Messitte's June 2026 walkthrough also mentions a hidden legacy Dialogue De-reverb module. That is distinct from De-reverb and from Dialogue Isolate's Reverb control. This workflow uses the documented current tools rather than depending on an unspecified hidden command. The Dialogue Isolate guide covers its voice, noise and reverb controls.
Use the rebuilt Breath Control deliberately
The RX 12 Breath Control manual describes three algorithms: Real-time, Offline and Classic. Classic uses the RX 11 algorithm; Offline uses more resources. For Real-time or Offline detection, the manual recommends longer selections or clips, preferably at least three seconds. Give the detector a phrase rather than only a tiny inhale.
Gain reduces detected breaths by the chosen amount. Target attenuates breaths toward the selected level, leaving quieter ones alone. Start gently and adjust Sensitivity by listening for mistaken consonants and whispered syllables; no fixed number suits every singer.
Natural preserves ambience during reduction; Gated removes it. Natural is a useful first comparison when room continuity matters. That switch is unavailable in Classic. Use the Editor's Listen ear button to check what is being removed, then switch it off. Do not confuse the Level control with a promise that every breath has been detected correctly.
Keep inhales that support the phrase. Check the result again through the intended vocal compression, because it can bring breaths forward. For one missed or exaggerated event, a local gain edit may be enough; protect its transitions instead of raising whole-take detection.
De-ess without changing the singer's diction
Judge sibilance after major repairs and again during mixing. The De-ess manual distinguishes Classic broadband attenuation from Spectral processing of the active high frequencies. Cutoff is the lower detection boundary, not a high-pass filter for the voice.
Compare several S, SH and F sounds, sustained vowels and the next word. Stop before diction turns lispy or the top end stays suppressed after the sibilant. If only two syllables are harsh, use local treatment or later automation. If Breath Control and De-ess make each other's artifacts more obvious, compare the opposite order from the same checkpoint.
Repair Assistant or a manual chain?
The current Repair Assistant manual explicitly assigns isolated sung vocals to Music mode. Voice is for spoken word and dialogue. In the Editor, choose Music, select representative vocal audio and click Learn. This analysis needs the performance and its defects, not a noise-only sample.
Compare the suggestion and bypass sections that do not help, including tone changes you would rather make during mixing. Music mode has a different set of controls from Voice; do not expect the assistant to reproduce every stage in this article. For the plug-in, click Learn before playing the relevant audio and let analysis finish.
Standard and Advanced support manual Module Chains. Transferring Repair Assistant's suggestion with Open Module Chain is Advanced-only. The Repair Assistant workflow explains that distinction. Save useful settings, but recheck learned profiles, selections and each enabled module on the next take.
Final vocal-cleanup quality gate
Compare each approved stage against its input at similar perceived loudness. If the vocal becomes dull, inspect click removal, denoising and de-essing; if low notes thin, inspect plosive treatment, hum filters and any high-pass filtering. Watery or pumping phrases call for less broad processing or a better noise profile. Undo the suspect stage before judging a replacement.
Use Preview and Compare where supported. If you rendered a short test, undo it before applying the chosen settings over a larger selection. Keep wanted vocals out of the removed-signal output, turn that monitor off, and apply each approved stage once. Check transitions, room continuity, breaths, consonants, sustained notes and output peaks.
Use File > Export to write a new WAV or AIFF filename at the intended delivery settings. Save overwrites an open WAV or AIFF. An RX Document keeps the editing history but opens only in RX, so it is not the audio file to hand back to the DAW.
Place the export at the original position in a session copy and check timing, duration and channels. Avoid leaving the same cleanup plug-ins active on an already repaired file. Listen through the entire returned vocal and mix, including quiet entrances and the last decay, before delivery. The Music & Vocals hub links the related cleanup and separation workflows.
Frequently asked questions
What order should I clean vocals in RX?
Use clipping and local faults, plosives and clicks, confirmed tonal noise, then broader noise or room treatment as a starting order. Balance breaths and sibilance afterward, changing their order if the comparison sounds better. Skip unnecessary stages.
Does this workflow work in RX Elements?
Elements supplies six plug-ins and no standalone Audio Editor. Use its available effects in your host; the full Editor workflow and dedicated Mouth De-click, De-plosive and Breath Control require Standard or Advanced.
Which Voice De-noise mode should I use for singing?
Select Optimize for Music to protect sustained sung notes. Compare gentle reduction first; use a representative noise-only sample if you choose Manual learning.
Should I use Mouth De-click on the whole vocal?
Test a phrase first. A light whole-take pass can help widespread clicks, but local repairs may preserve more detail. Use the Editor's Listen ear button to check removed material and disable it before rendering.
Should I remove every breath?
No. Retain breaths that support phrasing. In RX 12 Breath Control, compare Natural when ambience should remain, and use local gain edits for individual misses rather than escalating the entire take.
Which Repair Assistant mode is for sung vocals?
Music mode is the current manual's choice for isolated sung vocals. Voice is intended for spoken dialogue. Learn from representative performance audio, then compare and adjust the suggestion.
Why does RX make my vocal sound watery?
Excessive denoising, separation or stacked processing can cause that sound. Compare with the stage input, reduce processing or correct the noise profile, and repair remaining isolated defects locally.
Is vocal cleanup the same as removing vocals from a song?
No. This workflow repairs an isolated vocal recording. Music Rebalance separates a vocal from a finished mix, which can leave bleed and separation artifacts that require their own checks.



