Psychoacoustic Masking
Psychoacoustic Masking Models for Audio Perception · Also known as: masking, temporal masking, frequency masking, auditory masking
Psychoacoustic masking describes how the human auditory system suppresses the perception of weak sounds in the presence of stronger sounds. Formalized by Eberhard Zwicker in the 1960s, masking is a fundamental phenomenon in hearing and the basis for perceptual audio coding (MP3, AAC, OPUS). Masking occurs both in frequency (spectral masking) and time (temporal masking), and understanding these effects enables efficient audio compression and realistic sound design.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Apply psychoacoustic masking models in audio compression (MP3, AAC, OPUS), audio enhancement and noise reduction, hearing aid design, speech intelligibility assessment, and virtual audio rendering. Masking is essential for perceptual codec design where file size and bandwidth matter. Masking is less relevant when audio fidelity is paramount (studio mixing) or when perceptual metrics are not required.
Strengths & limitations
- Well-grounded in auditory physiology; masking thresholds correlate with measurements of basilar membrane vibration and cochlear processing.
- Enables dramatic audio compression (10:1 or higher) without perceptible quality loss by exploiting human hearing limitations rather than throwing away data.
- Applies across a broad range of listening conditions and acoustic environments; masking curves are fairly consistent across listeners and contexts.
- Simple and efficient computational implementation; masking models fit into real-time audio codecs and hearing aids with minimal CPU overhead.
- Established standards (ISO/MPEG) and decades of refinement; masking models are mature and reliable in commercial audio systems.
- Masking thresholds vary substantially across individuals; hearing loss, age, and individual differences reduce predictability in some listeners.
- Masking models assume linear auditory processing; nonlinear effects (loudness adaptation, distortion products) are not fully captured.
- Temporal masking is less well-characterized than frequency masking; models are simpler and less accurate for time-domain phenomena (transients, musical artifacts).
- Depends critically on signal statistics; masking thresholds change with spectrum, temporal structure, and signal type. Models may not generalize across all signals.
- Individual variability and context effects (attention, fatigue, prior exposure) modulate masking thresholds; fixed models cannot capture these variations.
Frequently asked
What is the difference between frequency masking and temporal masking?
Frequency (spectral) masking occurs when one frequency range masks perception of sounds in nearby frequencies (e.g., a loud low-frequency bass masks hearing of midrange instruments). Temporal masking occurs when a loud sound suppresses perception of softer sounds before it (pre-masking, ~5 ms) or after it (post-masking, ~100 ms). Both contribute to overall masking, but temporal masking is harder to model and less well-understood.
What is critical bandwidth and why does it matter?
Critical bandwidth is the width of the frequency region within which masking is strongest; it represents the effective 'tuning' of the cochlea (inner ear). Below 500 Hz, critical bandwidth is roughly 100 Hz; above 2 kHz, it widens. Audio codecs use critical bandwidth to group frequency bins for masking analysis and bit allocation—frequencies within a critical band are often treated as a unit.
How much can audio be compressed using masking?
Typical compression ratios are 10–20:1 for speech and 8–12:1 for music without noticeable quality loss (MP3 at 128 kbps, AAC at 96 kbps). Aggressive compression (40:1, 64 kbps) introduces audible artifacts (pre-echo, loss of transient detail, tonal artifacts). The achievable ratio depends on signal content, masking model quality, and listener hearing acuity.
Can hearing-impaired listeners perceive sounds that normal listeners find masked?
Yes. Hearing loss can shift masking thresholds. Some hearing-impaired listeners have elevated thresholds (need louder signals) or different frequency dependence. Additionally, recruitment (loudness growth) can make masked signals perceptible at higher levels. Standard masking models assume normal hearing; customization is needed for hearing-impaired applications.
What artifacts result from improper masking threshold estimation?
Under-estimating masking (allocating too few bits) causes audible quantization noise, pre-echo (artifacts before transients), and temporal smearing. Over-estimating masking (allocating too many bits) wastes space but ensures high quality. The most common artifacts are tonal buzzing (quantization of low-amplitude components) and loss of spatial detail in multi-channel audio.
Sources
- Zwicker, E., & Scharf, B. (1965). Psychoacoustics: Facts and Models. Springer-Verlag. ISBN: 978-3540631644
- Moore, B. C. J. (2012). An Introduction to the Psychology of Hearing (6th ed.). Academic Press. ISBN: 978-0123914232
- Johnston, J. D. (1988). Transform coding of audio signals using perceptual noise criteria. IEEE Journal on Selected Areas in Communications, 6(2), 314–323. DOI: 10.1109/49.608 ↗
How to cite this page
ScholarGate. (2026, June 3). Psychoacoustic Masking Models for Audio Perception. ScholarGate. https://scholargate.app/en/acoustics/psychoacoustic-masking
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Bark and Mel ScalesAcoustics↔ compare
- Cepstral AnalysisAcoustics↔ compare
- FxLMS Active Noise ControlAcoustics↔ compare
- Linear Predictive CodingAcoustics↔ compare
- Speech IntelligibilityAcoustics↔ compare