5  Cepstral Peak Prominence (CPP and CPPS)

Note

Living draft. This chapter reflects the sources cited below and will be revised as further sources are incorporated. Found an error? Use the Report an issue link in the sidebar.

5.1 Definition

Cepstral peak prominence (CPP) quantifies how strongly the harmonic structure of a voice signal emerges from its background. The cepstrum used in voice analysis is obtained by transforming the waveform into a spectrum, taking the logarithm, and transforming again; its horizontal axis is a time-like dimension called quefrency (Murton et al. 2020). The periodic harmonic peaks of the spectrum collapse into a single dominant cepstral peak at the quefrency of the voice period. The height of that peak above a regression line through the whole cepstrum is the CPP, reported in decibels (Murton et al. 2020). The amplitude of the dominant cepstral peak (the first rahmonic) reflects the strength with which the fundamental frequency emerges from competing background frequencies (Maryn et al. 2009).

Lower CPP values are associated with greater dysphonia severity (Murton et al. 2020). Two properties distinguish CPP from the classical perturbation measures. First, it can be extracted from connected speech as well as sustained vowels. Second, it does not require direct computation of the fundamental frequency, so no pitch detection or tracking is needed — which makes cepstral measures robust in dysphonic and particularly in breathy voices, where period-by-period analysis breaks down (Murton et al. 2020; Barsties v. Latoszek et al. 2018). The 2018 ASHA instrumental-assessment guidance recommended CPP as a general measure of dysphonia, replacing jitter, shimmer, and harmonics-to-noise ratio in that role (Murton et al. 2020).

5.2 Computation

The waveform is Fourier-transformed into a spectrum, the logarithm of the spectrum is taken, and a second (inverse) Fourier transform carries the result into the cepstral domain (Murton et al. 2020). A linear regression line relating quefrency to cepstral magnitude is fitted through the cepstrum, and CPP is the difference in amplitude between the cepstral peak and the point on the regression line directly below it (Maryn et al. 2009; Barsties v. Latoszek et al. 2018).

CPP versus CPPS. Averaging the cepstrum first across time and then across quefrency yields a smoothed cepstrum; the same peak-minus-regression-line measurement on the smoothed cepstrum is the smoothed cepstral peak prominence, CPPS (Maryn et al. 2009). Hillenbrand and colleagues developed SpeechTool, a program that derives the unsmoothed cepstrum; a later modification of the algorithm, averaging across time and then across quefrency, produced the smoothed cepstrum (Maryn et al. 2009).

Software implementations. Three computation algorithms dominate the clinical literature: Hillenbrand and Houde’s CPPS algorithm, Praat’s CPPS algorithm, and the CPP computation in Analysis of Dysphonia in Speech and Voice (ADSV, PENTAX Medical) (Murton et al. 2020). Their parameter choices differ in ways that matter:

  • ADSV applies a voicing activity detector, and frames with negative CPP (cepstral peak below the regression line) are excluded from analysis (Murton et al. 2020).
  • Praat computes CPPS from a PowerCepstrogram and includes no voicing activity detection. The reference configuration used by Murton et al. (2020): 60-Hz pitch floor, 2-ms time step, 5-kHz maximum frequency, pre-emphasis from 50 Hz, time averaging window 0.01 s, quefrency averaging window 0.001 s, peak search range 60–330 Hz, parabolic interpolation, straight tilt line over the quefrency range 0.001–0 s, robust fit.

Because of these differences, values from different programs live on different scales and are not interchangeable (see Confounds and cautions).

5.3 What it captures

CPP behaves as a general index of overall dysphonia severity. In the Massachusetts Eye and Ear Infirmary (MEEI) data of Murton et al. (2020), CPP values below task- and software-specific cutoffs indicated the presence of a voice disorder with up to 94.5% accuracy, and ADSV CPP estimated mean listener ratings of overall severity with r² from .5 to .74 depending on task (Praat CPPS: r² from .38 to .72). Regression lines allow a severity estimate from CPP; a Praat CPPS of 10 dB on sustained /a/ predicts a CAPE-V overall severity of approximately 54 (Murton et al. 2020).

Meta-analytically, CPP and CPPS are the best acoustic predictors of breathiness, and the only acoustic markers that reached the validity criterion for breathiness in both sustained vowels and continuous speech (Barsties v. Latoszek et al. 2018). Pooled effect sizes:

Table 5.1: Pooled meta-analytic correlations with perceptual ratings (Barsties v. Latoszek et al. 2018).
Perceptual dimension Task Measure Weighted r N
Breathiness sustained vowel CPP .66 468
Breathiness continuous speech CPP .73 98
Breathiness sustained vowel CPPS .64 133
Breathiness continuous speech CPPS .63 242
Roughness continuous speech CPPS .60 217

For roughness, CPPS reached the acceptability criterion only in continuous speech, where it was the only adequate predictor; CPPS does not, however, seem to distinguish the subtypes roughness and breathiness from each other well (Barsties v. Latoszek et al. 2018). Although the cepstral measures were originally designed for breathiness, they proved to be the most accurate acoustic measures of overall dysphonia severity (Barsties v. Latoszek et al. 2018).

In severely aperiodic and alaryngeal voices, where perturbation measures fail for lack of a reliably detectable fundamental period, CPP remains computable: in tracheoesophageal speakers it was the strongest single acoustic correlate of overall voice quality (rs = .78, about 60% of variance explained), and a two-factor regression combining CPP with the height of the second spectral harmonic reached rs = .87 with listener rankings (Maryn et al. 2009). That multivariate strategy is the same one the AVQI applies to laryngeal voice.

5.4 Normative data

CPP values and cutoffs belong to the pipeline that produced them: the program, its version and settings, and the speech material. Each row below states its pipeline and its compatibility class; see Pipeline Dependence and Reference Values for the classes. Compare a measured value only with a row whose pipeline matches.

Table 5.2: Diagnostic cutoffs (voice disorder vs. typical voice). MEEI rows: Murton et al. (2020). Second-sentence rows: Sauder et al. (2017). Sustained /i/: Buckley et al. (2023).
Population Task Pipeline Cutoff Accuracy Class
MEEI, English; 295 patients, 50 controls Sustained /a/ ADSV 3.4.2 CPP, default settings 11.46 dB 79.4% pipeline-compatible
MEEI, English; 291 patients, 50 controls Rainbow Passage, first 12 s ADSV 3.4.2 CPP, default settings 6.11 dB 87.7% pipeline-compatible
MEEI, English; 295 patients, 50 controls Sustained /a/ Praat 6.0.40 CPPS, Watts et al. settings 14.45 dB 77.4% pipeline-compatible
MEEI, English; 291 patients, 50 controls Rainbow Passage, first 12 s Praat 6.0.40 CPPS, Watts et al. settings 9.33 dB 94.5% pipeline-compatible
English; 100 voice-disordered, 70 controls Rainbow Passage, second sentence Praat 6.0.17 CPPS, factory settings 19.10 dB AUC 0.91 pipeline-compatible
— Sustained /i/ — No published cutoff in the sources reviewed — —

Other reported cutoffs are listed below for context only. Their pipelines are incomplete where they are reported: no software version, or values read from another paper’s summary table. They are not used as cutoffs here.

Table 5.3: Reported cutoffs without a complete pipeline, shown as history. Normal vs. dysphonic except Heman-Ackah 2003 (mild vs. severe). Sauder row: Sauder et al. (2017); other rows as tabulated by Murton et al. (2020).
Study Language Pipeline Sustained vowel Running speech Class
Sauder et al. 2017 English ADSV CPPS, factory settings; ADSV version not stated — 5.53 dB (second sentence of the Rainbow Passage; AUC 0.81) unclassified
Heman-Ackah et al. 2003 English CPPS (Hillenbrand) 10 dB 5 dB unclassified (as tabulated by Murton et al.)
Heman-Ackah et al. 2014 English CPPS (Hillenbrand) — 4.0 dB unclassified (as tabulated by Murton et al.)
Yu et al. 2018 Korean ADSV CPP 12 dB 7 dB unclassified (as tabulated by Murton et al.)
Delgado-Hernández et al. 2019 Spanish Praat CPPS (AVQI configuration) 13.96 dB 8.37 dB unclassified (as tabulated by Murton et al.)

Buckley et al. (2023) measured 150 speakers without voice disorders, aged 18–91 years, and proposed normative lower limits at 2 SD below the group mean. Praat CPPS used the Watts et al. settings (version 6.0.50); ADSV used default settings except a CPP threshold of 1 dB and vocalic event detection.

Table 5.4: Normative lower limits for typical speakers (Buckley et al. 2023). The ADSV rows are shown for comparison with the Praat rows; the ADSV version is not stated, so they are not used as reference limits.
Task Pipeline Males Females Class
Sustained /a/, middle 1 s Praat 6.0.50 CPPS, Watts et al. settings 11.72 dB 11.05 dB pipeline-compatible
Sustained /i/, middle 1 s Praat 6.0.50 CPPS, Watts et al. settings 12.01 dB 10.37 dB pipeline-compatible
Rainbow Passage, sentences 2–3 Praat 6.0.50 CPPS, Watts et al. settings 6.40 dB 6.49 dB pipeline-compatible
Sustained /a/, middle 1 s ADSV CPP; version not stated 8.86 dB 8.09 dB unclassified
Sustained /i/, middle 1 s ADSV CPP; version not stated 6.68 dB 4.65 dB unclassified
Rainbow Passage, sentences 2–3 ADSV CPP; version not stated 5.40 dB 5.04 dB unclassified

In that sample, 56% of females and 65.3% of males fell below the 9.33-dB Praat cutoff for the Rainbow Passage, and all fell below the 19.10-dB Praat cutoff (Buckley et al. 2023).

Thresholds depend on software, settings and task. Different algorithms produce CPP values in different ranges, and thresholds are lower for continuous speech than for sustained vowels (Murton et al. 2020). On the same samples, ADSV and Praat values correlate highly (r = 0.88), but their absolute values cannot be compared directly (Sauder et al. 2017).

5.5 Confounds and cautions

Software non-interchangeability. On the same MEEI recordings, the sustained-vowel cutoff was 11.46 dB for ADSV and 14.45 dB for Praat. The algorithms can also diverge qualitatively: one speaker measured 11.7 dB in Praat CPPS but 3.5 dB in ADSV CPP, with phonation showing irregular, widely spaced pulses perceived as strain and vocal fry (Murton et al. 2020).

Task effects. CPP thresholds are consistently lower for continuous speech than for sustained vowels. CPP varies widely between different sentences; one possible explanation is their differing amounts of voicing. Comparisons within or between speakers must therefore use the same speech material. The results suggest that including unvoiced frames in the average can artificially lower the CPP of consonant-heavy material (Murton et al. 2020).

Aphonia and voicing detection. Where aphonic segments are included, a low average CPP meaningfully reflects the aphonic quality — but very accurate voicing detection can paradoxically produce a higher-than-expected CPP in intermittent aphonia, because only the periodic voiced fragments enter the average (Murton et al. 2020).

No established upper bound. Murton et al. (2020) note that it is conceivable that some voice disorders lead to abnormally high CPP, which a single threshold would not take into account. They cite Awan and Awan (2020), who indicated that rough voices with a strong subharmonic component may exhibit high CPP values.

Vowel, loudness, sex, and age. Awan et al. (2012), as cited by Murton et al. (2020), found that low vowels such as /a/ tended to have higher CPP than high vowels such as /i/, and that CPP increases significantly with loudness. Murton et al. attribute the loudness effect to increased glottal closure and reduced perturbation, not to a change in underlying dysphonia. Male speakers tended to have higher CPP than female speakers, possibly through loudness. The MEEI controls were 22–59 years old; separate norms for older adults may be needed (Murton et al. 2020).

Cutoff vicinity. Published cutoffs are one possible estimate; values near a cutoff deserve further consideration rather than a binary reading (Murton et al. 2020).

Recording conditions. Hardware, microphone placement, environmental noise, and software are known to affect perturbation measures; their impact on CPP and CPPS remains unclear and needs further investigation (Barsties v. Latoszek et al. 2018).

5.6 Validation

The MEEI cutoff study (Murton et al. 2020) and the roughness/breathiness meta-analysis (Barsties v. Latoszek et al. 2018) anchor the evidence summarized above. The meta-analytic sustained-vowel breathiness effect sizes for CPP and CPPS pooled heterogeneous r-values; the authors retained these measures because they had the highest study counts and sample sizes (Barsties v. Latoszek et al. 2018). The tracheoesophageal findings of Maryn et al. (2009) extend the measure’s validity to alaryngeal voice.

5.7 Compute it in PhonaLab

Open PhonaLab →

PhonaLab is developed by the author of this reference; see Competing interests.

Upload a sustained vowel or a reading passage and compute CPPS on your own recording, with the software and task always reported alongside the value.

Barsties v. Latoszek, Ben, Youri Maryn, Ellen Gerrits, and Marc De Bodt. 2018. “A Meta-Analysis: Acoustic Measurement of Roughness and Breathiness.” Journal of Speech, Language, and Hearing Research 61 (2): 298–323. https://doi.org/10.1044/2017_JSLHR-S-16-0188.
Buckley, Daniel P., Defne Abur, and Cara E. Stepp. 2023. “Normative Values of Cepstral Peak Prominence Measures in Typical Speakers by Sex, Speech Stimuli, and Software Type Across the Life Span.” American Journal of Speech-Language Pathology 32 (4): 1565–77. https://doi.org/10.1044/2023_AJSLP-22-00264.
Maryn, Youri, Catherine Dick, Caroline Vandenbruaene, Tom Vauterin, and Tinne Jacobs. 2009. “Spectral, Cepstral, and Multivariate Exploration of Tracheoesophageal Voice Quality in Continuous Speech and Sustained Vowels.” The Laryngoscope 119 (12): 2384–94. https://doi.org/10.1002/lary.20620.
Murton, Olivia, Robert Hillman, and Daryush Mehta. 2020. “Cepstral Peak Prominence Values for Clinical Voice Evaluation.” American Journal of Speech-Language Pathology 29 (3): 1596–607. https://doi.org/10.1044/2020_AJSLP-20-00001.
Sauder, Cara, Michelle Bretl, and Tanya Eadie. 2017. “Predicting Voice Disorder Status from Smoothed Measures of Cepstral Peak Prominence Using Praat and Analysis of Dysphonia in Speech and Voice (ADSV).” Journal of Voice 31 (5): 557–66. https://doi.org/10.1016/j.jvoice.2017.01.006.