Process / pipelineLinguisticsAcoustic phoneticsPipeline

Acoustic Phonetic Analysis

Also known as: Acoustic Analysis of Speech, Speech Acoustic Measurement, Acoustic Speech Analysis

Acoustic phonetic analysis is the empirical measurement workflow at the heart of experimental phonetics: it records speech, segments and labels the signal, and extracts quantitative acoustic parameters — the waveform, the spectrogram, fundamental frequency (F0), the formants, intensity, segment duration, and voice onset time (VOT). These measurements are interpreted through the source-filter theory of speech production, which models the output sound as a glottal source spectrum shaped by the transfer function of the vocal tract, turning the audible speech stream into reproducible numbers that can be compared, modelled, and related to articulation.

Key highlights

  • Yields objective, numerical, reproducible measurements that replace impressionistic auditory judgments with verifiable data.
  • Grounded in the well-understood physics of the source-filter theory, giving acoustic parameters a principled link to articulation.
  • Non-invasive and inexpensive: a single recording supports measurement of pitch, formants, duration, intensity, and timing at once.
  • Provides a common quantitative currency that connects phonetics to sociolinguistics, clinical work, forensics, and speech technology.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use acoustic phonetic analysis whenever you need objective, reproducible, quantitative evidence about how speech sounds are physically realized — vowel quality, consonant place and voicing, pitch and intonation, timing, and voice quality. It is the methodological backbone of experimental phonetics, sociophonetics, clinical speech assessment, forensic speaker comparison, and speech-technology research. It is less appropriate when the research question is about higher-level meaning, pragmatics, or perception that the acoustic signal alone cannot settle, or when recordings are too noisy or uncontrolled to yield trustworthy measurements.

Strengths & limitations

Strengths
  • Yields objective, numerical, reproducible measurements that replace impressionistic auditory judgments with verifiable data.
  • Grounded in the well-understood physics of the source-filter theory, giving acoustic parameters a principled link to articulation.
  • Non-invasive and inexpensive: a single recording supports measurement of pitch, formants, duration, intensity, and timing at once.
  • Provides a common quantitative currency that connects phonetics to sociolinguistics, clinical work, forensics, and speech technology.
Limitations
  • Measurements are only as good as the recording: noise, reverberation, clipping, and inconsistent microphone distance corrupt parameters.
  • The acoustic-to-articulatory mapping is many-to-one, so the same acoustic output can arise from different articulatory configurations.
  • Automatic estimators — especially LPC formant tracking and pitch tracking — make errors that require manual checking and correction.
  • Acoustic parameters vary with speaker anatomy (vocal-tract length, sex, age), so raw values are not directly comparable across speakers without normalization.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

How is acoustic phonetic analysis different from acoustic phonetics as a field?

Acoustic phonetics is the subdiscipline that studies the physical, acoustic properties of speech sounds and the theory (notably the source-filter model) that explains them. Acoustic phonetic analysis is the empirical measurement workflow practitioners actually carry out — recording, segmenting, and measuring parameters such as F0, formants, duration, intensity, and VOT from real signals. In short, the field supplies the theory and questions; the analysis is the hands-on procedure that produces the numbers used to answer them.

Why is the source-filter theory central to interpreting the measurements?

The source-filter theory says the output speech spectrum is the product of an independent source (glottal vibration or turbulence) and a filter (the vocal-tract resonances). This separation is what lets acoustic measurements be read as articulatory evidence: F0 and voice quality reflect the source, while the formant frequencies reflect the filter — the shape and length of the vocal tract. Without this model the formants and pitch would be just numbers; with it, they become a window onto how the sound was produced.

What sampling rate and recording quality do I need?

By the Nyquist theorem the sampling rate must be at least twice the highest frequency you want to measure. A rate of 16 kHz captures the main formant range for vowels; 22.05 kHz or 44.1 kHz is preferable when fricatives or fine spectral detail matter. Equally important are a quiet, low-reverberation environment, a flat-response microphone at a fixed distance, and recording levels that avoid clipping — because no amount of later processing can recover information the recording failed to capture.

Sources

  1. 1.
    Johnson, K. (2012). Acoustic and Auditory Phonetics (3rd ed.). Wiley-Blackwell.
    ISBN 9781405194662
  2. 2.
    Ladefoged, P., & Johnson, K. (2014). A Course in Phonetics (7th ed.). Cengage.
    ISBN 9781285463407
  3. 3.
    Boersma, P., & Weenink, D. (2023). Praat: Doing phonetics by computer [Computer program].

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Acoustic Phonetic Analysis. ScholarGate. https://scholargate.app/linguistics/acoustic-phonetic-analysis