On this page
Clean audiobook narration in iZotope RX 12 chapter by chapter, but manage it as one book. Lock the manuscript and delivery specification first, group files by recording session, repair clipping and hum before broad denoise, remove only distracting mouth noise and plosives, then level and measure every complete chapter. The steps below use the standalone RX Audio Editor. Keep the untouched recordings and lossless edit masters separate from the final upload files.
Do not master to a remembered “audiobook preset.” As checked on September 8, 2026, ACX requires −23 to −18 dB RMS, peaks no higher than −3 dB and a noise floor below −60 dB RMS. Spotify's current guide recommends a broader −24 to −14 dB RMS range. Those are different destinations, and a file that passes one sheet is not automatically ready for the other.
Audiobook cleanup at a glance
| Stage | RX decision | Approval question |
|---|---|---|
| Ingest | Duplicate lossless source; group by session | Can every edit be reversed? |
| Content edit | Remove takes, mistakes and long interruptions | Does the reading still follow the locked manuscript? |
| Hard repair | De-clip, De-hum, local Spectral Repair | Is the actual defect gone without changing the voice? |
| Noise | Voice/Spectral De-noise or Dialogue Isolate | Is room tone quieter but still continuous? |
| Mouth and air | Mouth De-click, De-plosive, selective breath edits | Are consonants and performance intact? |
| Consistency | Gain/level control, matched transitions | Do adjacent chapters sound like the same book? |
| Delivery | Measure, encode, reopen and audition | Does every final file meet this distributor's current sheet? |
Start with the destination, not the noise-reduction module
Before opening RX, obtain the current platform specification and the production's editorial sheet. Confirm chapter divisions, opening and closing credits, narrator names, required sample, file naming, mono or stereo delivery, codec, sample rate, level metric, peak ceiling, noise floor and boundary spacing. Mark which rules are requirements and which are recommendations.
The current ACX Audio Submission Requirements apply per file. The Spotify for Authors upload guide accepts MP3, WAV and FLAC and asks for clear chapter titles. Recheck the chosen page immediately before export: delivery rules are operational facts, not permanent audio theory.
ACX and Spotify audiobook specifications are not the same
| Check | ACX current requirement | Spotify current guide |
|---|---|---|
| Average level | −23 to −18 dB RMS per file | Recommended −24 to −14 dB RMS |
| Peak | No higher than −3 dB | No numeric peak ceiling stated in the linked guide; check the project sheet |
| Noise floor | Below −60 dB RMS | Recommended below −60 dB |
| Head/tail | Recommended 1–5 seconds of room tone; not over 5 seconds | Recommended 0.5–1 second head, 1–5 seconds tail |
| Audio file | 192 kbps+ CBR MP3, 44.1 kHz | MP3: 192 kbps+, 44.1 kHz, CBR preferred; WAV: 44.1 kHz/16-bit; FLAC accepted |
| Channels | All mono or all stereo | Do not mix mono and stereo |
Spotify's values come from the Metadata & Asset Guide linked by its current upload page. The PDF is marked November 2024; it lists those RMS, noise and spacing values under “Recommended qualities.” Use it alongside any newer delivery sheet supplied for your project. ACX's page also requires consistent tone, level, noise, spacing and pronunciation, and calls out plosives, pops, mouse clicks, excessive mouth noise and outtakes. Numeric compliance is only half the job.
Protect the source and build a session map
Create an untouched archive of the original recordings. Edit duplicates in a lossless format and keep the original sample rate and channel layout until the delivery stage. Use File > Save RX Document to retain the RX session and its edit history. Ordinary Save on a WAV or AIFF overwrites that file; export a separately named delivery copy. Do not repeatedly encode MP3 while editing; create compressed upload copies only after the approved master exists.
Map every chapter to narrator, microphone, preamp/interface, room, recording date, gain, distance and known interruption. A narrator moving closer after lunch can matter more than the chapter number. Group files by matching conditions. This session map decides where one RX chain is reusable and where it is not.
The safe-save and versioning details are covered in our RX import, export and saving guide. Keep an edit master, a processed master and the final delivery copy as separate assets.
Edit content before mastering the book
Remove false starts, repeated sentences, direction, chair moves and interruptions against the locked manuscript. Preserve deliberate pauses, sentence rhythm and character performance. Use crossfades and matching room tone where a cut needs a fill. Inserting digital silence into an otherwise audible room bed can make the background switch conspicuously on and off.
Do not polish a two-hour chapter and then discover that twelve minutes are outtakes. Content editing first reduces processing time and prevents later cuts from exposing different noise treatments. Listen across each splice on headphones and speakers. The correct edit should be invisible even before denoise.
Build one diagnostic sample per recording condition
Choose a passage long enough to include clean room tone, ordinary narration, a quiet phrase, a loud phrase, breaths, mouth clicks, sibilance, plosives and the worst background event. Include edit boundaries. This sample is a stress test for the entire session, not a flattering excerpt.
Save an RX Document before committing a repair, then change one process at a time. Preview alone does not create a saved edit. Compare from the same pre-roll at matched loudness. Our Preview, Compare and History workflow keeps candidate chains auditable. If a setting helps normal speech but damages the quiet phrase, it is not ready for the chapter.
Use an audiobook-safe repair order
- Content edits: approved takes, timing and room-tone fills.
- Hard defects: clipping, hum, isolated bumps and digital clicks.
- Broad noise/reverb: only the stable or changing problem actually present.
- Mouth and air: lip smacks, plosives and exceptional breaths.
- Tone and consistency: restrained EQ or level correction when required.
- Delivery level: final whole-file RMS and peak checks, plus noise-floor checks on actual room tone from each file.
This order is not ritual. Clipping changes peaks, hum contaminates a noise profile, and aggressive denoise can alter click detection. Fixing the strongest causal defect first makes later modules work less. The broader RX Module Chain guide explains why chain order must follow the signal.
Repair clipping, hum and isolated bumps first
If flat-topped peaks or audible crunch are present, follow the De-clip workflow before level control. Use a representative clipped passage, inspect the waveform and keep an untouched comparison. De-clip estimates missing peak shape; it cannot recover a pristine performance from severe overload.
For narrow electrical tones and harmonics, use the De-hum workflow. For a page hit, bump or single chair squeak, select only the event and use Spectral Repair or attenuation. Do not run a global repair because one noise occurs twice in a book.
Choose denoise from the noise behavior
Stable HVAC hiss or preamp noise is a Voice De-noise or Spectral De-noise problem. Changing traffic, crowd bleed or room reverberation may justify Dialogue Isolate. The detailed decision tree in Voice De-noise vs Spectral De-noise prevents the common mistake of reaching for the most powerful module first.
Use the least processing that removes distraction at normal listening level. A faint steady room is less tiring than watery consonants that move with every sentence. Audit quiet words, whispered dialogue, sibilants and long vowels. If the room tone collapses only while the narrator speaks, back off.
The current RX 12 Dialogue Isolate manual separates Voice, Noise and Reverb controls. Use that flexibility only when the recording has the corresponding problem; our Dialogue Isolate tutorial covers sensitivity and artifact checks in depth.
Treat room tone as continuity, not waste
Room tone carries the acoustic identity of a session. Keep labeled tone beds for every room and narrator setup. Use them under edits, at pickups and at chapter boundaries. When a pickup was recorded on another day, match level and tone locally; do not flatten the whole chapter to disguise one mismatch.
Measure noise floor in true room tone, not a digitally silent edit, a breath or the decaying end of a word. A meter can report an impressive number for inserted zeros while the narration still contains HVAC noise. Listen through headphones at a realistic gain and inspect the spectrogram for tonal or changing components.
Remove mouth clicks without erasing diction
Run Mouth De-click on the diagnostic sample, not the entire book by reflex. Start conservatively and enable the footer’s Listen ear icon during Preview to hear the removed signal. If you hear consonants, lip articulation that belongs to a character, or fragments of words, reduce Sensitivity and repair the remaining obvious clicks locally. Turn Listen off, audition normal speech again and leave it off before Render or export.
The official Mouth De-click manual describes detection and repair controls; the common module controls explain Listen. Our Mouth De-click guide covers Sensitivity, Frequency Skew and Click Widening. “No mouth sound at all” is not the target; uninterrupted listening is.
Handle plosives, breaths and sibilance selectively
Low-frequency P and B bursts belong in De-plosive. Process the affected syllable with enough surrounding audio to judge the result, then check that the word has not become thin. Run De-plosive before a high-pass filter: filtering first can interfere with its low-frequency detection. The RX De-plosive tutorial shows the local workflow.
Breaths are timing and performance. Shorten or lower the distracting ones, especially after a comp, but do not erase every inhale. A breathless narrator sounds edited and can make long passages uncomfortable. Treat sibilance only when it is genuinely harsh; check several S sounds because one global setting can darken the whole voice.
Use Trim Silence as a supervised edit
RX 12 Trim Silence can remove detected silent regions and provides Threshold, Silence, Post-roll and Crossfade controls. It ignores the current selection and processes the entire file. iZotope's current fast-cleanup guide demonstrates it for spoken-word work, and the RX 12 manual documents the controls. It is Advanced-only.
For audiobooks, global silence removal can destroy dramatic timing, clip quiet syllable tails and violate required head/tail room tone. Work on a duplicate chapter, run the proposed trim, then listen to the resulting file and inspect every removal boundary. Undo or restore the working copy if it cuts wanted material. Treat it as a supervised time edit, not automatic mastering. Never apply it after final spacing without rebuilding and rechecking the delivery boundaries.
Match chapters without sanding away the narrator
Choose a representative approved chapter as a tonal reference. Compare adjacent chapter openings at the same monitor level. Listen for changes in proximity, low-frequency buildup, brightness, room size and background texture. Repair session discontinuities with the smallest local move; broad EQ across a good chapter can create a second problem.
In Advanced, Leveler can smooth level changes; restrained manual gain adjustment is another option. Leveler aims at a K-weighted RMS target and may not hit it exactly. Its result still needs measurement against the delivery specification. A shouted scene and a whispered confession should not have identical short-term behavior. The book needs a stable listening experience while preserving the performance's intended dynamics.
Do not confuse RMS, LUFS, peak and noise floor
- RMS: an average-energy measurement used in the current ACX and Spotify audiobook sheets.
- LUFS: a loudness metric common in broadcast and podcast delivery; it is not a synonym for the stated ACX RMS range.
- Sample peak: the highest encoded sample value; ACX states its ceiling as peak level.
- True peak: an estimate of inter-sample peaks after reconstruction; useful for encoding safety but not interchangeable with every platform's wording.
- Noise floor: the unwanted background level in actual room tone, not inserted digital silence.
Open Waveform Statistics and select the complete exported chapter for Total RMS and peak readings. A loud paragraph does not represent the file average. Then select genuine room tone to assess the background; the minimum RMS found somewhere in a chapter is not automatically its noise floor. Check several room-tone passages if the background changes.
RX’s Calculate RMS using AES-17 preference changes the RMS reference by 3 dB compared with its square-wave reference. Record the setting, use a consistent reference across chapters and compare the actual encoded file with the distributor’s validator. Do not fix an unexplained meter disagreement by changing gain blindly. Loudness Control targets LUFS/LKFS and true peak; entering an audiobook RMS number into its LUFS field does not establish compliance.
Batch only chapters that have already earned the same chain
After a full pilot chapter passes listening and measurement, lock its Module Chain and apply it only to files from the same recording condition. The RX 12 Batch Processor manual documents processing, naming and output. Our safe Batch Processor workflow covers pilot runs and rollback.
Write processed files to a new destination. Do not batch pickups, guest narrators or a noisy remote chapter through the main narrator's chain. After the run, inspect the quietest and loudest chapters and a heavily edited file first to catch a bad chain early. Then complete listening QC on every finished chapter. Check that each output exists, contains the intended processing and has the expected duration; a successful batch dialog alone does not approve the book.
Build files, credits and chapter order deliberately
Both platforms limit each audio file to 120 minutes. ACX requires one chapter or section per file, with separate opening and closing credits and a secondary header when a long section is split. Spotify also calls for separate chapter and introduction files; for an unchaptered book, its guide specifies 30–120-minute segments. Follow the destination’s chapter structure and naming convention.
Check opening credits word for word against the platform sheet. Spotify's current guide calls for the title plus author and narrator names. Confirm multi-narrator roles, chapter labels, sequence numbers and sample boundaries. A technically clean master with chapter 12 and 13 reversed is a failed delivery. For ACX, include a retail sample no longer than five minutes from the book and follow its content restrictions. Spotify’s linked guide also caps samples at five minutes and excludes music, credits and explicit content. Its closing-credit passages conflict over whether those credits are mandatory; resolve that point in the project’s delivery sheet instead of silently treating them as optional.
Ten-step RX 12 audiobook workflow
- Lock the manuscript and delivery sheet. Confirm chapters, credits, names, format, level, peak, noise and spacing rules.
- Protect the source recordings. Keep untouched lossless masters and work on versioned duplicates.
- Group files by recording session. Separate narrator, microphone, room, gain and noise conditions.
- Build a diagnostic sample. Include room tone, normal and quiet speech, loud lines, breaths, mouth noise and the worst defect.
- Repair hard defects first. Correct clipping, hum, isolated bumps and digital spikes before broad processing.
- Reduce stable or changing noise conservatively. Match Voice/Spectral De-noise or Dialogue Isolate to the real noise behavior.
- Treat mouth clicks, plosives and breaths selectively. Protect consonants, natural inhalation and performance timing.
- Run only proven batches. Approve a full pilot chapter, process matching files to a new folder and inspect the actual outputs.
- Level and measure each chapter. Match the book, check full-file RMS and peaks, and measure actual room tone separately.
- Export and verify encoded files. Reopen every final file, repeat its measurements, check names, order, credits, spacing, channels and codec, and complete listening QC.
Final audiobook quality-control checklist
- Read chapter starts, endings and credits against the locked manuscript.
- Confirm every narrator is credited and chapter files are in order.
- Listen through every complete chapter against the manuscript, including edits, pickups and room-tone fills.
- Check quiet words, loud lines, consonants, breaths, plosives and mouth clicks.
- Compare adjacent chapters for tone, room, level and spacing.
- Measure full-file RMS and peaks; check the noise floor in real room-tone passages from every file.
- Confirm all files use the required codec, sample rate and one channel format.
- Reopen and audition the final encoded upload files.
- Run the ACX Audio Lab when delivering through ACX; it checks eight metrics including RMS, peak, bitrate and spacing. ACX’s audio quality FAQ says it does not detect noise-floor or editing problems.
- Retain the source, edit master, processed master and delivery copy separately.
The RX Workflows hub connects this book-length process with Batch Processor, Module Chains, RX Connect and DAW round trips. For shorter mixed episodes, use the separate podcast cleanup workflow.
What not to do
- Do not use ACX numbers as a universal audiobook specification.
- Do not batch chapters recorded in different rooms through one preset.
- Do not fill edits with digital silence when real room tone is available.
- Do not remove every breath, pause or mouth movement.
- Do not let Trim Silence decide performance timing unattended.
- Do not measure a short selection and assume the whole chapter passes.
- Do not overwrite original recordings or lossless edit masters.
- Do not approve from meters without listening to the encoded files.
Frequently asked questions
What is the best iZotope RX chain for audiobook cleanup?
There is no safe universal chain. Edit content first, repair clipping or hum if present, choose denoise from the noise behavior, treat mouth clicks and plosives selectively, then level and measure the complete chapter. Prove the chain on a representative sample before batching.
What RMS level does ACX currently require?
ACX's current submission page requires each file to measure between −23 and −18 dB RMS, with peaks no higher than −3 dB and a noise floor below −60 dB RMS. Recheck the official page before delivery.
What level does Spotify recommend for audiobooks?
Spotify's current Audiobook Metadata & Asset Guide recommends −24 to −14 dB RMS and noise below −60 dB. That recommendation is broader than ACX's range and should not be substituted for another distributor's sheet.
Should I remove every breath from audiobook narration?
No. Breaths carry pacing and performance. Lower or shorten distracting inhales, especially around edits, but retain natural breathing unless the producer or rights holder requests a specific style.
Can Mouth De-click run across a whole chapter?
Only after a representative sample and pilot chapter prove that it removes clicks without taking consonants or articulation. Use the Listen ear icon during Preview to inspect removed audio, then switch it off and check normal speech before Render. Use conservative settings and spot-repair the remaining events.
Is RX Trim Silence safe for audiobooks?
It can speed supervised editing in RX 12 Advanced, but an unattended global pass can damage quiet syllables, dramatic pauses and required head/tail room tone. It always processes the whole file, regardless of the selection. Test on a duplicate and listen to the rendered result, checking every boundary.
Can I batch-process all audiobook chapters?
Batch only files from the same recording condition after one full pilot chapter passes listening and measurement. Put pickups, guest narrators and changed rooms in separate groups, and always write to a new destination.
If the audiobook passes RMS and peak checks, is it ready?
Not necessarily. The book can still contain bad edits, wrong credits, chapter-order errors, inconsistent tone, mouth noise or processing artifacts. Numeric validation and complete listening QC are separate gates.



