5 Cepstral Peak Prominence (CPP and CPPS)
Living draft. This chapter reflects the sources cited below and will be revised as further sources are incorporated. Found an error? Use the Report an issue link in the sidebar.
5.1 Definition
Cepstral peak prominence (CPP) quantifies how strongly the harmonic structure of a voice signal emerges from its background. The cepstrum used in voice analysis is obtained by transforming the waveform into a spectrum, taking the logarithm, and transforming again; its horizontal axis is a time-like dimension called quefrency (Murton et al. 2020). The periodic harmonic peaks of the spectrum collapse into a single dominant cepstral peak at the quefrency of the voice period. The height of that peak above a regression line through the whole cepstrum is the CPP, reported in decibels (Murton et al. 2020). The amplitude of the dominant cepstral peak (the first rahmonic) reflects the strength with which the fundamental frequency emerges from competing background frequencies (Maryn et al. 2009).
Lower CPP values are associated with greater dysphonia severity (Murton et al. 2020). Two properties distinguish CPP from the classical perturbation measures. First, it can be extracted from connected speech as well as sustained vowels. Second, it does not require direct computation of the fundamental frequency, so no pitch detection or tracking is needed — which makes cepstral measures robust in dysphonic and particularly in breathy voices, where period-by-period analysis breaks down (Murton et al. 2020; Barsties v. Latoszek et al. 2018). The 2018 ASHA instrumental-assessment guidance recommended CPP as a general measure of dysphonia, replacing jitter, shimmer, and harmonics-to-noise ratio in that role (Murton et al. 2020).
5.2 Computation
The waveform is Fourier-transformed into a spectrum, the logarithm of the spectrum is taken, and a second (inverse) Fourier transform carries the result into the cepstral domain (Murton et al. 2020). A linear regression line relating quefrency to cepstral magnitude is fitted through the cepstrum, and CPP is the difference in amplitude between the cepstral peak and the point on the regression line directly below it (Maryn et al. 2009; Barsties v. Latoszek et al. 2018).
CPP versus CPPS. Averaging the cepstrum first across time and then across quefrency yields a smoothed cepstrum; the same peak-minus-regression-line measurement on the smoothed cepstrum is the smoothed cepstral peak prominence, CPPS (Maryn et al. 2009). Hillenbrand and colleagues developed SpeechTool, a program that derives the unsmoothed cepstrum; a later modification of the algorithm, averaging across time and then across quefrency, produced the smoothed cepstrum (Maryn et al. 2009).
Software implementations. Three computation algorithms dominate the clinical literature: Hillenbrand and Houde’s CPPS algorithm, Praat’s CPPS algorithm, and the CPP computation in Analysis of Dysphonia in Speech and Voice (ADSV, PENTAX Medical) (Murton et al. 2020). Their parameter choices differ in ways that matter:
- ADSV applies a voicing activity detector, and frames with negative CPP (cepstral peak below the regression line) are excluded from analysis (Murton et al. 2020).
- Praat computes CPPS from a PowerCepstrogram and includes no voicing activity detection. The reference configuration used by Murton et al. (2020): 60-Hz pitch floor, 2-ms time step, 5-kHz maximum frequency, pre-emphasis from 50 Hz, time averaging window 0.01 s, quefrency averaging window 0.001 s, peak search range 60–330 Hz, parabolic interpolation, straight tilt line over the quefrency range 0.001–0 s, robust fit.
Because of these differences, values from different programs live on different scales and are not interchangeable (see Confounds and cautions).
5.3 What it captures
CPP behaves as a general index of overall dysphonia severity. In the Massachusetts Eye and Ear Infirmary (MEEI) data of Murton et al. (2020), CPP values below task- and software-specific cutoffs indicated the presence of a voice disorder with up to 94.5% accuracy, and ADSV CPP estimated mean listener ratings of overall severity with r² from .5 to .74 depending on task (Praat CPPS: r² from .38 to .72). Regression lines allow a severity estimate from CPP; a Praat CPPS of 10 dB on sustained /a/ predicts a CAPE-V overall severity of approximately 54 (Murton et al. 2020).
Meta-analytically, CPP and CPPS are the best acoustic predictors of breathiness, and the only acoustic markers that reached the validity criterion for breathiness in both sustained vowels and continuous speech (Barsties v. Latoszek et al. 2018). Pooled effect sizes:
| Perceptual dimension | Task | Measure | Weighted r | N |
|---|---|---|---|---|
| Breathiness | sustained vowel | CPP | .66 | 468 |
| Breathiness | continuous speech | CPP | .73 | 98 |
| Breathiness | sustained vowel | CPPS | .64 | 133 |
| Breathiness | continuous speech | CPPS | .63 | 242 |
| Roughness | continuous speech | CPPS | .60 | 217 |
For roughness, CPPS reached the acceptability criterion only in continuous speech, where it was the only adequate predictor; CPPS does not, however, seem to distinguish the subtypes roughness and breathiness from each other well (Barsties v. Latoszek et al. 2018). Although the cepstral measures were originally designed for breathiness, they proved to be the most accurate acoustic measures of overall dysphonia severity (Barsties v. Latoszek et al. 2018).
In severely aperiodic and alaryngeal voices, where perturbation measures fail for lack of a reliably detectable fundamental period, CPP remains computable: in tracheoesophageal speakers it was the strongest single acoustic correlate of overall voice quality (rs = .78, about 60% of variance explained), and a two-factor regression combining CPP with the height of the second spectral harmonic reached rs = .87 with listener rankings (Maryn et al. 2009). That multivariate strategy is the same one the AVQI applies to laryngeal voice.
5.4 Normative data
CPP values and cutoffs belong to the pipeline that produced them: the program, its version and settings, and the speech material. Each row below states its pipeline and its compatibility class; see Pipeline Dependence and Reference Values for the classes. Compare a measured value only with a row whose pipeline matches.
| Population | Task | Pipeline | Cutoff | Accuracy | Class |
|---|---|---|---|---|---|
| MEEI, English; 295 patients, 50 controls | Sustained /a/ | ADSV 3.4.2 CPP, default settings | 11.46 dB | 79.4% | pipeline-compatible |
| MEEI, English; 291 patients, 50 controls | Rainbow Passage, first 12 s | ADSV 3.4.2 CPP, default settings | 6.11 dB | 87.7% | pipeline-compatible |
| MEEI, English; 295 patients, 50 controls | Sustained /a/ | Praat 6.0.40 CPPS, Watts et al. settings | 14.45 dB | 77.4% | pipeline-compatible |
| MEEI, English; 291 patients, 50 controls | Rainbow Passage, first 12 s | Praat 6.0.40 CPPS, Watts et al. settings | 9.33 dB | 94.5% | pipeline-compatible |
| English; 100 voice-disordered, 70 controls | Rainbow Passage, second sentence | Praat 6.0.17 CPPS, factory settings | 19.10 dB | AUC 0.91 | pipeline-compatible |
| — | Sustained /i/ | — | No published cutoff in the sources reviewed | — | — |
Other reported cutoffs are listed below for context only. Their pipelines are incomplete where they are reported: no software version, or values read from another paper’s summary table. They are not used as cutoffs here.
| Study | Language | Pipeline | Sustained vowel | Running speech | Class |
|---|---|---|---|---|---|
| Sauder et al. 2017 | English | ADSV CPPS, factory settings; ADSV version not stated | — | 5.53 dB (second sentence of the Rainbow Passage; AUC 0.81) | unclassified |
| Heman-Ackah et al. 2003 | English | CPPS (Hillenbrand) | 10 dB | 5 dB | unclassified (as tabulated by Murton et al.) |
| Heman-Ackah et al. 2014 | English | CPPS (Hillenbrand) | — | 4.0 dB | unclassified (as tabulated by Murton et al.) |
| Yu et al. 2018 | Korean | ADSV CPP | 12 dB | 7 dB | unclassified (as tabulated by Murton et al.) |
| Delgado-Hernández et al. 2019 | Spanish | Praat CPPS (AVQI configuration) | 13.96 dB | 8.37 dB | unclassified (as tabulated by Murton et al.) |
Buckley et al. (2023) measured 150 speakers without voice disorders, aged 18–91 years, and proposed normative lower limits at 2 SD below the group mean. Praat CPPS used the Watts et al. settings (version 6.0.50); ADSV used default settings except a CPP threshold of 1 dB and vocalic event detection.
| Task | Pipeline | Males | Females | Class |
|---|---|---|---|---|
| Sustained /a/, middle 1 s | Praat 6.0.50 CPPS, Watts et al. settings | 11.72 dB | 11.05 dB | pipeline-compatible |
| Sustained /i/, middle 1 s | Praat 6.0.50 CPPS, Watts et al. settings | 12.01 dB | 10.37 dB | pipeline-compatible |
| Rainbow Passage, sentences 2–3 | Praat 6.0.50 CPPS, Watts et al. settings | 6.40 dB | 6.49 dB | pipeline-compatible |
| Sustained /a/, middle 1 s | ADSV CPP; version not stated | 8.86 dB | 8.09 dB | unclassified |
| Sustained /i/, middle 1 s | ADSV CPP; version not stated | 6.68 dB | 4.65 dB | unclassified |
| Rainbow Passage, sentences 2–3 | ADSV CPP; version not stated | 5.40 dB | 5.04 dB | unclassified |
In that sample, 56% of females and 65.3% of males fell below the 9.33-dB Praat cutoff for the Rainbow Passage, and all fell below the 19.10-dB Praat cutoff (Buckley et al. 2023).
Thresholds depend on software, settings and task. Different algorithms produce CPP values in different ranges, and thresholds are lower for continuous speech than for sustained vowels (Murton et al. 2020). On the same samples, ADSV and Praat values correlate highly (r = 0.88), but their absolute values cannot be compared directly (Sauder et al. 2017).
5.5 Confounds and cautions
Software non-interchangeability. On the same MEEI recordings, the sustained-vowel cutoff was 11.46 dB for ADSV and 14.45 dB for Praat. The algorithms can also diverge qualitatively: one speaker measured 11.7 dB in Praat CPPS but 3.5 dB in ADSV CPP, with phonation showing irregular, widely spaced pulses perceived as strain and vocal fry (Murton et al. 2020).
Task effects. CPP thresholds are consistently lower for continuous speech than for sustained vowels. CPP varies widely between different sentences; one possible explanation is their differing amounts of voicing. Comparisons within or between speakers must therefore use the same speech material. The results suggest that including unvoiced frames in the average can artificially lower the CPP of consonant-heavy material (Murton et al. 2020).
Aphonia and voicing detection. Where aphonic segments are included, a low average CPP meaningfully reflects the aphonic quality — but very accurate voicing detection can paradoxically produce a higher-than-expected CPP in intermittent aphonia, because only the periodic voiced fragments enter the average (Murton et al. 2020).
No established upper bound. Murton et al. (2020) note that it is conceivable that some voice disorders lead to abnormally high CPP, which a single threshold would not take into account. They cite Awan and Awan (2020), who indicated that rough voices with a strong subharmonic component may exhibit high CPP values.
Vowel, loudness, sex, and age. Awan et al. (2012), as cited by Murton et al. (2020), found that low vowels such as /a/ tended to have higher CPP than high vowels such as /i/, and that CPP increases significantly with loudness. Murton et al. attribute the loudness effect to increased glottal closure and reduced perturbation, not to a change in underlying dysphonia. Male speakers tended to have higher CPP than female speakers, possibly through loudness. The MEEI controls were 22–59 years old; separate norms for older adults may be needed (Murton et al. 2020).
Cutoff vicinity. Published cutoffs are one possible estimate; values near a cutoff deserve further consideration rather than a binary reading (Murton et al. 2020).
Recording conditions. Hardware, microphone placement, environmental noise, and software are known to affect perturbation measures; their impact on CPP and CPPS remains unclear and needs further investigation (Barsties v. Latoszek et al. 2018).
5.6 Validation
The MEEI cutoff study (Murton et al. 2020) and the roughness/breathiness meta-analysis (Barsties v. Latoszek et al. 2018) anchor the evidence summarized above. The meta-analytic sustained-vowel breathiness effect sizes for CPP and CPPS pooled heterogeneous r-values; the authors retained these measures because they had the highest study counts and sample sizes (Barsties v. Latoszek et al. 2018). The tracheoesophageal findings of Maryn et al. (2009) extend the measure’s validity to alaryngeal voice.
5.7 Compute it in PhonaLab
PhonaLab is developed by the author of this reference; see Competing interests.
Upload a sustained vowel or a reading passage and compute CPPS on your own recording, with the software and task always reported alongside the value.