7 Acoustic Breathiness Index (ABI)
Living draft. This chapter reflects the sources cited below and will be revised as further sources are incorporated. Found an error? Use the Report an issue link in the sidebar.
7.1 Definition
The Acoustic Breathiness Index (ABI) is a multiparametric, nine-variable acoustic measure that quantifies the degree of perceived breathiness with a single score, computed from concatenated samples of continuous speech and the sustained vowel /a/ (Barsties v. Latoszek et al. 2021). Its perceptual target is breathiness: turbulent noise, excessively high in frequency, resulting from air leakage during incomplete glottal closure (Barsties v. Latoszek et al. 2021; Delgado Hernández et al. 2018). Both speech tasks contribute three seconds of voiced material each, a design chosen for ecological validity (Barsties v. Latoszek et al. 2021). The index was first published in 2017 for a Dutch-speaking population (Barsties v. Latoszek et al. 2017, 2021).
Where the AVQI targets overall voice quality, the ABI isolates one perceptual dimension. The two indices share the concatenation protocol and several components.
7.2 Computation
The ABI contains nine parameters (Delgado Hernández et al. 2018):
| Parameter | Abbreviation | What it measures |
|---|---|---|
| Smoothed cepstral peak prominence | CPPS | Height of the first rahmonic’s peak over the regression line through the smoothed cepstrum — see the CPPS chapter |
| Jitter local | Jit | Average difference between successive periods, divided by the average period |
| Glottal-to-noise excitation ratio | GNEmax-4500 Hz | Excitation due to vocal-fold oscillation versus excitation by turbulent noise; reliable even under strong amplitude and frequency perturbations |
| High-frequency noise | Hfno-6000 Hz | Relative level of high-frequency noise between the energy from 0–6 kHz and from 6–10 kHz |
| Harmonics-to-noise ratio of Dejonckere | HNR-D | Harmonic emergence of the spectral display in the 500–1500 Hz band |
| Amplitude difference of the first two harmonics | H1–H2 | Indirect measure of the relative length of the open phase of glottal oscillation; H1 is relatively high in breathy voices |
| Shimmer local dB | ShdB | Base-10 logarithm of the difference between amplitudes of successive periods, × 20 |
| Shimmer local | Shim | Absolute mean difference between amplitudes of successive periods, divided by the average amplitude |
| Period standard deviation | PSD | Variation in the standard deviation of periods |
The rescaled regression equation, as published in the development study (Barsties v. Latoszek et al. 2017):
\[ \begin{aligned} \text{ABI} = \bigl(5.0447730915 &- 0.172\,\text{CPPS} - 0.193\,\text{Jit} - 1.283\,\text{GNE}_{\max\text{-}4500}\\ &- 0.396\,\text{Hfno}_{6000} + 0.01\,\text{HNR-D} + 0.017\,(\text{H1–H2})\\ &+ 1.473\,\text{ShdB} - 0.088\,\text{Shim} - 68.295\,\text{PSD}\bigr) \times 2.9257400394 \end{aligned} \]
The German validation uses the same intercept (Barsties v. Latoszek et al. 2020); the Spanish and US English validations print 5.0447740915 (Delgado Hernández et al. 2018; Castillo-Allendes et al. 2023). The difference changes a score by about 3 × 10−6 and has no clinical effect.
The score ranges from 0 to 10; the higher the ABI, the more severe the breathiness (Barsties v. Latoszek et al. 2021).
The development paper provides its Praat script as supplementary data (Barsties v. Latoszek et al. 2017); this script is the reference implementation of the equation above. Later work refers to it as ABI script v. 01.01 (Stappenbeck et al. 2020).
Signal processing uses Praat with the customised ABI script (Barsties v. Latoszek et al. 2021). Analysis applies only to voiced segments of the continuous speech, extracted with an automated Praat detection script, with a three-second /a/ segment appended (Delgado Hernández et al. 2018). The continuous-speech part is standardised to a language-specific syllable count yielding three seconds of voiced speech. The Spanish validation set 33 syllables for both AVQIv3 and ABI; the 34-syllable (Dutch) and 30-syllable (Japanese) cut-offs it cites were established for AVQIv3 (Delgado Hernández et al. 2018).
7.3 What it captures
ABI scores correlate strongly with auditory-perceptual breathiness ratings: Spearman rs ranged from 0.746 to 0.890 across the six validation studies pooled in the meta-analysis (Barsties v. Latoszek et al. 2021). In Spanish, rs = 0.826, with ABI accounting for 68.2% of the variance in mean breathiness ratings (Delgado Hernández et al. 2018); in the original Dutch study, rs = 0.840 (Delgado Hernández et al. 2018).
Unlike the CSID and AVQI, which assess overall voice quality, the ABI is particularly suitable for evaluating breathy voices — for example benign vocal-fold lesions dominantly characterised by breathiness (nodules of medium or large size), paralysis or paresis of the recurrent laryngeal nerve, and vocal-fold bowing associated with presbyphonia (Barsties v. Latoszek et al. 2021); acute laryngitis also belongs to this group (Delgado Hernández et al. 2018). Neither age and gender nor roughness significantly affects ABI in natural voices, and the index signals therapy-related voice quality changes with high sensitivity (Barsties v. Latoszek et al. 2021).
Independent meta-analysis of single measures supports the component choice: Hfno, H1–H2, HNR of Dejonckere, CPP, and CPPS all sit among the strongest single correlates of perceived breathiness, and only 12 of 85 acoustic measures predicted breathiness on sustained vowels with a weighted mean r of at least .60 (Barsties v. Latoszek et al. 2018).
7.4 Normative data
ABI validation evidence applies to Praat with the customised ABI script (Barsties v. Latoszek et al. 2021). Each row states its pipeline and compatibility class; see Pipeline Dependence and Reference Values.
| Population / language | Pipeline | Threshold | Sensitivity | Specificity | AUC | Class |
|---|---|---|---|---|---|---|
| Dutch, development study (970 dysphonic, 88 healthy) | Praat 5.3.57, custom ABI script | 3.44 | 82.4% | 92.9% | 0.948 | language-conditional |
| Spanish (136 dysphonic, 47 healthy) | Praat 6.0.22, ABI script | 3.40 | 73.7% | 95.4% | 0.921 | language-conditional |
| German (175 dysphonic, 43 healthy) | Praat 5.3.57, ABI script | 3.42 | 72% | 95% | 0.91 | language-conditional |
| Brazilian Portuguese, counting 1–11 (40 dysphonic, 13 healthy; 17 syllables) | Praat 6.0.06, ABI script | 2.38 | 95.2% | 81.8% | 0.924 | language-conditional |
| Brazilian Portuguese, reading text (same 53 speakers; 32 syllables) | Praat 6.0.06, ABI script | 3.13 | 88.1% | 90.9% | 0.929 | language-conditional |
| US English (148 with voice disorders, 49 non-clinic-seeking; 22 syllables) | VOXplot 2.0.0; Praat 6.3.06 for extraction | 2.35 | 84% | 81% | 0.89 | language-conditional |
| Finnish (108 dysphonic, 87 non-dysphonic; 31 syllables) | VOXplot 2.0.0; recorded in Praat 6.2.23 | 2.68 | 80% | 80% | 0.886 | language-conditional |
| Pooled, six studies (context only) | Praat + ABI script; heterogeneous hardware | none | 0.84 (95% CI 0.83–0.85) | 0.92 (95% CI 0.89–0.94) | 0.94 (SROC) | meta-analytic |
The six validation studies pooled by the meta-analysis cover Dutch, German, Spanish, Brazilian Portuguese, Korean, and Japanese — four language groups — over 467 vocally healthy and 3136 voice-disordered subjects (Barsties v. Latoszek et al. 2021). The meta-analysis also derived a weighted threshold of 3.40 across its six studies (sensitivity 0.86, specificity 0.90) (Barsties v. Latoszek et al. 2021). This reference does not use it as a cutoff: a pooled value mixes languages, recording set-ups and Praat versions, and pooled sensitivity was highly heterogeneous (I² = 89.7%) (Barsties v. Latoszek et al. 2021). Use the row for the patient’s language and task.
Table 7.3 lists the threshold each study reported.
| Language | Study | Healthy (N) | Disordered (P) | Concurrent validity (\(r_s\)) | Threshold | Class |
|---|---|---|---|---|---|---|
| Dutch | Barsties v. Latoszek et al. | 88 | 954 | 0.840 | 3.44 | language-conditional |
| Spanish | Delgado Hernández et al. | 47 | 136 | 0.826 | 3.40 | language-conditional |
| Japanese | Hosokawa et al. | 55 | 288 | 0.890 | 3.44 | language-conditional |
| Brazilian Portuguese | Englert et al. | 37 | 113 | 0.746 | 2.94 | language-conditional |
| German | Barsties v. Latoszek et al. | 43 | 175 | 0.850 | 3.42 | language-conditional |
| Korean | Kim et al. | 197 | 1470 | 0.870 | 3.69 | language-conditional |
Thresholds range from 2.94 (Brazilian Portuguese) to 3.69 (Korean); see Language dependence of the threshold below. The later US English validation, not part of the meta-analysis, reports 2.35 (LR+ 4.29, LR− 0.2) and \(r_s\) = 0.77 between ABI and perceived breathiness (Castillo-Allendes et al. 2023). Its Methods define the ROC groups by the overall grade G for both indices, while its figure caption describes the ABI curve as separating breathy from nonbreathy voices, so the perceptual reference for the ABI threshold is not fully clear. The Finnish validation, also later than the meta-analysis, reports 2.68 (LR+ 4.00, LR− 0.25) and \(r_s\) = 0.823 with perceived breathiness, with nonbreathy voices defined as mean B below 0.5 (Kankare and Laukkanen 2023). The Brazilian Portuguese value comes from a study the meta-analysis cites as a one-page 2019 item in the International Archives of Otorhinolaryngology (vol. 23, p. 106); its inter-rater reliability for breathiness was below the recommended level (Barsties v. Latoszek et al. 2021). The later full Brazilian Portuguese validation (Englert et al. 2021) covers the AVQI only.
A second Brazilian Portuguese study, of 53 speakers, reports two ABI thresholds from the same people: 2.38 for counting 1–11 and 3.13 for the first sentence of a read text. Mean ABI scores were higher for reading (Englert et al. 2020). The task alone moves the threshold by 0.75. The study had only 13 vocally healthy speakers, and its authors describe a full validation as still in process (Englert et al. 2020). The 3.13 that appears for Brazilian Portuguese in the Finnish paper’s summary table (Kankare and Laukkanen 2023) is this reading-text value. That table has errors, and the Englert et al. table prints LR+ values that cannot be correct, so neither is used here for likelihood ratios. For Brazilian Portuguese, use the row that matches the task: 2.94 comes from the larger sample, but the meta-analysis does not state its task.
7.5 Confounds and cautions
Language dependence of the threshold. Continuous speech is part of the index, so inter-language phonetic differences may influence the outcome, and cross-validation studies are needed to establish the ABI’s validity for each language. The meta-analysis nevertheless found ABI relatively robust to phonetic inter-language differences such as stress timing, syllable timing, and intonation (Barsties v. Latoszek et al. 2021).
The Praat version matters. With the same ABI script and the same 218 German recordings, seven Praat versions formed three clusters. A CPPS bug introduced in Praat 6.0.44 and fixed in 6.0.47 cut sensitivity at the German threshold of 3.42 from 72% to 9% with version 6.0.46. Across the unaffected versions, mean differences were negligible, but single recordings differed by up to about 3 points (Stappenbeck et al. 2020). Use the Praat version of the validation study for the language in question, or check a new version against an earlier one on test recordings (Stappenbeck et al. 2020).
Software requirement. The validation evidence applies to signal processing in Praat with the customised ABI script; the meta-analysis excluded studies not using them (Barsties v. Latoszek et al. 2021).
Equal task proportion. ABI results based on unequal proportions of continuous speech and sustained vowel were treated as a risk of bias (Barsties v. Latoszek et al. 2021).
Hardware and recording conditions. Each validation study has unique acoustic settings and hardware, limiting comparability between studies (Barsties v. Latoszek et al. 2021). The Spanish validation recorded in a soundproof booth with a head-mounted condenser microphone at 44.1 kHz / 16 bit (Delgado Hernández et al. 2018).
Evidence base. The meta-analysis rests on six studies. Pooled sensitivity showed high heterogeneity (I² = 89.7%, P < .001), while specificity heterogeneity was low (I² = 45.5). The Brazilian Portuguese study reached only low inter-rater reliability (Fleiss κ < 0.41), with a sample skewed toward normal and mild breathiness (43% normal, 41% mild, 11% moderate, 5% severe) (Barsties v. Latoszek et al. 2021).
Not a stand-alone screen. Not all voice disorders are detected by ABI — that cannot be expected from a measure of one perceptual dimension; higher diagnostic accuracy requires combining measures of different dimensions (Barsties v. Latoszek et al. 2021). Further investigations of external and internal validity remain necessary (Delgado Hernández et al. 2018).
7.6 Validation
The development study established the Dutch threshold of 3.44 with AUC 0.948 (Barsties v. Latoszek et al. 2017; Delgado Hernández et al. 2018). The Spanish validation reported a threshold of 3.40 with AUC 0.921 (Delgado Hernández et al. 2018). The 2021 meta-analysis pooled six validation studies spanning more than 3600 voice samples and confirmed the ABI as a robust, valid objective measure of breathiness, with pooled sensitivity 0.84 and specificity 0.92 (Barsties v. Latoszek et al. 2021).
7.7 Compute it in PhonaLab
PhonaLab is developed by the author of this reference; see Competing interests.
Upload a sustained vowel and a continuous speech sample to compute ABI on your own recording, alongside the AVQI computed from the same material.