iZotope RX World
Workflow

Clean Up Podcast Audio with iZotope RX 12

By Brandon Hayes Published Sep 4, 2026 Revised Sep 7, 2026 13 min read

Short answer

A repeatable RX 12 podcast workflow that diagnoses each speaker, repairs only the defects present and keeps loudness compliance for the final mixed episode.

Podcast microphone and noisy speech waveform passing through staged audio cleanup to a controlled final level
Editorial visualization of a podcast cleanup workflow; not an iZotope interface screenshot.

A reliable iZotope RX podcast cleanup starts with a defect map, not a preset. Keep each speaker separate, repair clipping and intrusive noise before tonal mixing, treat clicks and plosives only where they exist, then level the dialogue and measure loudness on the finished episode. The goal is not a perfectly silent spectrogram. It is speech that stays intelligible, consistent and human for an entire listen.

RX 12 can speed up this work, but one chain cannot understand that the host has air-conditioner hum while the remote guest has room echo and codec damage. Build a small tested chain for each recurring source, keep an exception queue for odd noises, and leave delivery loudness until music, ads and transitions are in place.

This workflow uses the standalone RX 12 Audio Editor in Standard or Advanced. RX Elements contains plug-ins only. Standard includes the core noise and click tools, Module Chain, Batch Processor and Loudness Control; Trim Silence and Leveler require Advanced.

Podcast cleanup at a glance

What you hearFirst RX routeDo not confuse it with
Flat-topped peaks, crunchy syllablesDe-clipOrdinary loudness or compression
Steady electrical tone and harmonicsDe-humBroadband room noise
Fan or stable room noise behind speechVoice De-noise or Spectral De-noiseChanging traffic, crowd or reflections
Changing noise or obvious room reverbDialogue IsolateA learned stationary-noise profile
Low-frequency microphone “p” burstDe-plosiveA mouth click or digital click
Saliva clicks and lip smacksMouth De-clickPlosives and consonant attacks
Single cough, chair squeak or dropped objectLocal selection / Spectral RepairA global denoise pass
Host and guest drift in levelClip gain or LevelerFinal LUFS compliance
Finished episode misses delivery targetLoudness Control and true-peak checkPer-speaker leveling

That last distinction prevents a common failure. Leveling makes a conversation easier to follow from phrase to phrase. Loudness control adjusts a completed program toward a measured specification. If you force every raw track to a release target first, the final mix changes it again.

Protect the recording before processing

Keep the received WAV, AIFF or original recorder files untouched. Make working copies, name them by speaker and preserve the original sample rate and channel layout. If you were sent one stereo call recording, confirm whether each side contains a different speaker before converting anything to mono. A reversible session lets you return to clean evidence when a repair turns brittle.

In RX, Save overwrites an uncompressed WAV or AIFF. Use Save As for a separate working copy, and Save RX Document for a checkpoint containing the original and edit history. An .rxdoc opens only in RX; export audio for your DAW. When returning separate tracks, preserve their common start and full duration so they still line up.

Separate-track recording is worth protecting because each voice carries a different room, microphone distance and noise floor. A threshold that helps the host can chew through a softer remote guest. If you only have a flattened mix, make smaller time selections whenever the dominant speaker or background changes.

Edit content before polishing throwaways

Choose takes, remove false starts and build the conversation before spending twenty minutes restoring an answer that will be cut. iZotope's broader podcast production guide also separates content editing from later cleanup and mixing. This order saves time and gives you the true transitions that need room tone.

There is one practical exception: when severe clipping or noise makes an editorial decision impossible, repair a short diagnostic copy first. Do not process two hours merely to discover that only twelve minutes survive the edit. For long monologues or audiobooks, RX 12's Trim Silence can accelerate obvious gaps, but threshold-based deletion deserves a full review.

Trim Silence with room tone in mind

iZotope's May 2026 RX 12 podcast cleanup walkthrough begins with Trim Silence. This Advanced-only module can shorten repeatable pauses in solo narration. The Trim Silence manual warns that it processes the entire file and ignores your selection. Test a separately exported short copy first, then review the full processed recording.

Set Threshold below the quietest speech you must keep. Post-roll preserves the beginning of each detected silent section for word endings and room tails; Silence controls how much silence remains after that post-roll. Crossfade softens the joins. Audition every cut around laughter, interruptions and off-mic acknowledgements: a level threshold cannot judge whether a quiet answer matters.

Do not trim synchronized speaker tracks independently: deleting different pauses shifts their timing relative to one another and to video. Make linked structural edits in your DAW, keeping the common timeline intact. Use standalone Trim Silence only on a copy whose timing you are free to change, such as solo narration. Preserve believable room tone instead of forcing every gap to digital silence.

Make a defect map for each speaker

Listen once before opening modules. Drop markers for five categories: continuous noise, changing background, overload/distortion, mouth and breath events, and isolated interruptions. Also mark the cleanest noise-only passage, the quietest essential phrase and the loudest laugh. Those three references keep denoise and leveling decisions honest.

The current iZotope audio cleanup guide makes the same first move: identify whether you are hearing steady noise, plosives, clipping or distortion before choosing a tool. “Noisy” is not a diagnosis. Hum, room reflection, codec warble and a chair squeak require different processing.

Repair clipping and hard failures first

Use De-clip only where the waveform and sound support a clipping diagnosis. Flat peaks and a hard edge on vowels are better evidence than a loud meter alone. Repairing clipped samples before compression, leveling or limiting gives later processors a less distorted signal to react to. Our RX De-clip workflow covers threshold placement, gain staging and before/after checks.

Set the De-clip threshold at the actual clipped level, which may be below 0 dBFS after an earlier gain change. Leave headroom for reconstructed peaks. A louder result is not proof that the distorted syllable was repaired.

Do not expect De-clip to restore a microphone that never captured the voice clearly. Phone codecs, packet loss and distance from the mic may leave missing detail that no peak reconstruction can recover. Repair dropouts and isolated glitches locally; if a word remains unintelligible, an editorial pickup or transparent acknowledgement is better than synthetic certainty.

Choose the right noise-reduction route

For spoken word, the useful first comparison is usually Voice De-noise versus Spectral De-noise. Voice De-noise is designed for efficient, zero-latency work and can follow a changing noise floor in Adaptive mode. For manual learning, turn Adaptive off, select only background noise and click Learn; iZotope recommends at least one second. Then select the speech you intend to process, rather than leaving only the noise sample selected.

Choose Optimize for Dialogue and start with Gentle when transparency matters. The RX 12 Voice De-noise manual says Surgical offers stronger reduction but can produce chirpy or watery artifacts. Stop increasing reduction when consonants become papery, room tone pumps between words or the guest sounds detached from the space.

Spectral De-noise gives more detailed control over a learned fan, hiss or broadband noise profile. Learn from noise without speech, then select the target audio. Its higher quality modes demand more processing and latency; Adaptive mode can also follow changing noise, with greater resource use. De-hum is more direct for a tonal mains component. For a brief keyboard strike or truck pass, test a local repair before increasing reduction across the entire interview.

Use Dialogue Isolate for changing noise and reflections

The current Dialogue Isolate manual describes separation of spoken dialogue from constant or non-stationary backgrounds such as hiss, crowds, traffic, footsteps and weather. It also separates room reflections, with independent Voice, Reverb and Noise gain controls. This makes it a strong candidate for a remote guest whose noise changes underneath speech.

Dialogue Isolate is a module and plug-in in Standard and Advanced. Raise Sensitivity slowly: higher values classify more sound as noise and reverb, but can reduce dialogue clarity. Advanced adds Quality choices and four-band control in both the module and plug-in. Best/Offline trades processing time for fewer artifacts and supports offline Editor processing or DAW renders. Test the worst sentence and the quietest sentence before rendering the whole interview.

For separate component editing, RX 12 Advanced can send dialogue, noise and reverb to Stems View. These are separated sound components, not the four frequency bands. Stem Split ignores the gain settings and creates stems at 0 dB; do not assume your noise attenuation carries over. If you only want the current gain balance applied, use Render. Our RX Stems View guide covers editing and export: All Stems export includes every stem, and Single file exports their sum. Check that you are delivering the cleaned voice, not adding the separated noise back.

Remove plosives before high-pass filtering

A microphone pop is a short low-frequency pressure event, not ordinary bass tone. RX's De-plosive documentation says detection operates from 20 to 80 Hz and can fail after a high-pass filter has removed that evidence. Run De-plosive before corrective high-pass EQ when a true plosive is present.

Increase Sensitivity only enough to catch the bursts. High Strength can reduce clarity, while high Sensitivity may classify more speech as plosive. Preview phrases containing legitimate low fundamentals and hard consonants. A pop that becomes a small, believable thump is usually safer than a surgically hollow first syllable.

The 20–80 Hz range describes detection. Frequency Limit sets the upper boundary of reduction, so it is a different control. Adjust it against the actual pop and the speaker's low voice tones, rather than treating 80 Hz as a universal repair cutoff.

Remove mouth clicks without erasing diction

The current Mouth De-click manual says the module detects clicks and lip smacks across long selections or individual events. Sensitivity determines how many events it catches; too much can affect plosives and damage the original signal. Click Widening extends the repair around sounds with a short decay.

Test a dense thirty-second passage rather than assuming a preset is safe. Listen to “t,” “k,” “p” and tongue-detail consonants at normal speed. A second light pass can sometimes reveal quieter clicks masked by louder ones, as iZotope notes, but compare it with one carefully adjusted pass. Two passes are not automatically gentler.

Treat breaths as performance, not garbage

Breaths carry pacing and emotion. Remove only the distracting inhalations: clipped gasps, close-mic blasts and breaths made unnaturally loud by later compression. The standalone RX Breath Control guide shows how to reduce rather than erase them.

RX 12 Advanced Leveler also includes a desired breath-level control: it reduces louder detected breaths while leaving breaths already below that target alone. First audition the raw breath against the sentence, then choose attenuation that preserves timing. Silence before every phrase makes a host sound assembled rather than present.

Repair coughs, chair squeaks and interruptions locally

Global chains earn their keep on repeated defects. Exceptions belong in local selections. Use Spectral Repair, attenuate, replace, short fades or a clean room-tone patch according to the event. Do not send the entire guest track through stronger denoise because a truck passed during one answer.

If a cough completely covers a word, a cleaner spectrogram cannot prove the missing speech was recovered. Check the other speaker's microphone or a backup recording; otherwise use a pickup or revise the edit. Do not turn an estimated repair into a confident claim about what was said.

Keep an exception queue while editing: timestamp, problem, chosen repair and approval state. The spectrogram makes short broadband impacts, whistles and tonal rings easier to distinguish. Our selection, Preview, Compare and History guide helps audition candidates from the same pre-roll and return to a known revision.

Build one Module Chain per recurring source

The RX 12 Module Chain runs several modules in series and lets you add, remove, reorder, bypass and save the chain. That is ideal for a recurring host recorded in the same treated room. Name the preset with the microphone, room and purpose—not “Perfect Podcast.”

A defensible starting chain might contain light De-hum, Voice De-noise, Mouth De-click and a final safety gain step, but only if the defect map calls for all four. Leave De-plosive before any high-pass filter. Keep variable, high-risk jobs such as strong Dialogue Isolate or global Spectral Repair out until a representative section has passed an A/B check.

The full Module Chain guide covers ordering and preset design. The important podcast rule is simpler: a chain is a saved hypothesis for one repeatable source, not a diagnosis for every guest.

Batch only homogeneous recordings

Open Window > Batch Processor in RX Standard or Advanced. The Batch Processor can process multiple files with a custom Module Chain and displays channel count, bit depth, sample rate and duration. Its presets can retain chain and output choices, making it useful for same-room narration or chapters captured with one locked setup.

Before committing the batch, run one typical file, one quiet file and one difficult file. Exclude guests, pickups and remote calls whose noise or level differs. Choose a new output folder and a clear suffix; never overwrite received originals. The RX batch-processing guide covers the setup. Spot checks qualify the settings; listen through every finished file before approving the batch. Keep independent silence trimming out of synchronized multitrack jobs.

In a July 2022 r/podcasting discussion about RX 8 and RX 9, editors described batch cleanup before DAW editing, while one contributor adjusted the chain for different guests. That is a historical workflow example, not an RX 12 quality test or a reason to skip listening.

Level speakers before final loudness

Balance each speaker after repair so the listener does not ride the volume control. Manual clip gain is often best for obvious phrase changes. RX 12 Advanced Leveler can automate time-variable gain, creating a visible Clip Gain envelope that you can edit. The Leveler manual says it aims toward K-weighted RMS and may not hit the target exactly.

That is not a defect; Leveler smooths variation rather than certifying final LUFS. Higher Responsiveness values work across words or phrases more smoothly, while lower values react more aggressively. Check laughs, interjections and breaths so the algorithm does not pull noise forward. Applying another module after Leveler bakes its Clip Gain envelope into the audio; use Undo History or an RX Document checkpoint if you need the earlier editable state.

Mix before you set delivery loudness

Add music, ads, transitions and ambience, then measure the complete episode. RX Loudness Control applies a fixed gain to meet a chosen standard; it can add post-limiting when the true-peak specification requires it. Its integrated control uses LKFS, which the manual identifies as the same unit as LUFS.

For a concrete reference, Apple Podcasts recommends approximately -16 LKFS, within ±1 LU, with true peaks no higher than -1 dBFS, measured before encoding. This is Apple's recommendation, not a universal rule for every network. Its page does not prescribe a separate -19 LUFS mono target. Follow a client's delivery sheet when it specifies different requirements, and record the target, tolerance, true-peak ceiling and format.

In Loudness Control, select the entire finished episode, enter Integrated loudness, Tolerance and True peak for the delivery requirement, and check the applicable gating settings before Render. Do not select a preset by its -16 label alone: RX's AES streaming preset allows ±2 LU, while Apple recommends ±1. Leveler's K-weighted RMS target is a separate setting and cannot replace this measurement.

For an ordinary Apple Podcasts RSS feed, delivery audio must be MP3 or AAC. Subscriber audio uploaded through Apple Podcasts Connect has separate format and channel rules. Keep a lossless master; RX exports WAV, AIFF, FLAC, OGG and MP3, so use a suitable external encoder if delivering AAC. Reopen the actual encoded file and measure it again: encoding can change peaks. RX's MP3 Prevent Clipping option targets 0 dBTP, so it does not by itself verify a -1 true-peak ceiling.

Loudness compliance does not prove a good mix. After rendering, listen at normal volume and on a small speaker. A compliant episode can still contain pumping room tone, buried guests, piercing sibilance or music that masks speech.

Step-by-step RX 12 podcast workflow

  1. Protect and organize the recordings. Keep untouched source files, work on duplicates and preserve separate speaker tracks.
  2. Build a defect map. Mark clipping, hum, changing noise, reflections, plosives, mouth clicks, breaths and isolated interruptions.
  3. Finish structural edits. Choose takes and remove unwanted conversation before polishing material that will not survive.
  4. Repair damaged peaks and gross defects. De-clip genuine overload and repair isolated failures before later dynamics.
  5. Reduce hum, noise and reflections. Choose De-hum, Voice De-noise, Dialogue Isolate or Spectral De-noise according to the diagnosed defect.
  6. Treat plosives and mouth clicks. Run De-plosive before high-pass filtering and audition Mouth De-click against consonants.
  7. Repair exceptions by hand. Use local selections for coughs, squeaks and intermittent events.
  8. Level each speaker. Use clip gain or Leveler while preserving natural phrase dynamics.
  9. Mix the complete episode. Add music and transitions, retain believable room tone and check every edit in context.
  10. Measure and export. Apply the actual delivery specification, check true peaks, listen through the export and retain a lossless master.

Save a checkpoint after each stage. In the Audio Editor, bypassing a module does not undo processing already rendered into the file: return through Undo History or reopen the checkpoint for a fair comparison. The Preview Bypass control compares processing while previewing. Compare the same sentence at matched listening levels before accepting another stage.

Quality control before release

  • Listen to the first and last word around every structural edit.
  • Check the quietest guest answer and the loudest laugh.
  • Solo each speaker for artifacts, then judge the final decision in the mix.
  • Confirm room tone does not pump, loop obviously or vanish between sentences.
  • Verify “p,” “t,” “k” and sibilant consonants survived automated cleanup.
  • Check lip sync if the podcast also has video.
  • Measure integrated loudness and true peak on the final exported program.
  • Listen through the encoded delivery file, not only the lossless session.
  • Retain the untouched source, cleaned stems, final lossless master and delivery copy.

In a June 2024 discussion about a reverberant interview, the original poster worried about artifacts from stacking noise-reduction plug-ins. Replies proposed several different tools and combinations; there was no agreed winning chain. Treat this as a source-specific caution, not a current RX 12 comparison. Judge your episode against its unchanged source.

What not to do

  • Do not run one preset across every speaker and remote guest.
  • Do not overwrite the original recordings.
  • Do not remove every breath or make every pause digitally silent.
  • Do not high-pass before De-plosive when low-frequency pop evidence is needed.
  • Do not mistake louder output for successful denoising.
  • Do not use Leveler as proof of final LUFS compliance.
  • Do not set delivery loudness before the music, ads and transitions are mixed.
  • Do not batch guests whose recording conditions differ.

The RX Workflows hub connects this podcast process with DAW round trips, batch jobs, Module Chain design and other repeatable production routes.

Frequently asked questions

What order should I clean podcast audio in RX?

Protect the source, finish the structural edit, repair clipping and hard failures, reduce diagnosed noise, treat plosives and mouth clicks, fix exceptions locally, level each speaker, mix the episode, then set and verify final loudness.

Should I edit before or after noise reduction?

Edit first by default so you do not restore material that will be removed. Make a short diagnostic repair first only when severe damage prevents you from choosing takes or understanding speech.

Which RX module removes podcast background noise?

Use Voice De-noise or Spectral De-noise for stable noise, De-hum for tonal hum, and Dialogue Isolate for changing noise or strong room reflections. Diagnose the sound before choosing the module.

Is Dialogue Isolate better than Voice De-noise?

Neither is universally better. Voice De-noise is efficient for spoken-word noise floors; Dialogue Isolate can separate speech, changing noise and reverb. Compare them on the quietest essential sentence and keep the more natural result.

How do I remove mouth clicks from a podcast?

Test Mouth De-click on a dense passage, raise Sensitivity only until the distracting clicks are caught, and audition hard consonants. Use Click Widening for decaying lip smacks and local selections for stubborn exceptions.

Can I batch-process podcast audio in RX?

Yes, in RX Standard or Advanced, when files come from a consistent source and representative files have passed review. Exclude different microphones, guests, pickups and remote calls. Write to a new output folder and listen through every result; do not independently trim silence from synchronized tracks.

How loud should a podcast be?

Apple Podcasts recommends approximately -16 LKFS, within ±1 LU, with true peaks no higher than -1 dBFS before encoding. Follow your network or client specification when different. Measure the complete mix rather than each raw speaker track, then recheck the encoded export.

Does RX replace good podcast recording technique?

No. RX can reduce captured defects, but it cannot recover information the microphone never recorded. Mic position, room treatment, gain staging, separate tracks and a local backup remain the strongest cleanup tools.