6 Acoustic Voice Quality Index (AVQI)
Living draft. This chapter reflects the sources cited below and will be revised as further sources are incorporated. Found an error? Use the Report an issue link in the sidebar.
6.1 Definition
The Acoustic Voice Quality Index (AVQI) is an objective method to quantify the severity of overall voice quality in concatenated continuous speech and sustained phonation segments (Barsties and Maryn 2016). It is a multivariate construct based on linear regression that combines six acoustic markers into a single score for the whole voice sample (Barsties and Maryn 2016). That score represents the overall voice quality — the hoarseness level — of the sample (Delgado Hernández et al. 2018).
The index was proposed by Maryn, Corthals, et al. (2010) and was one of the first measures to combine both speech types (Delgado Hernández et al. 2018). The rationale for the combination: sustained vowels are easily elicited and less affected by articulation and dialect, while continuous speech is more representative of a patient’s daily voice (Delgado Hernández et al. 2018).
6.2 Computation
6.2.1 Speech material and concatenation
The subject sustains the vowel [a:] for longer than 3 seconds, and a 3-second mid-vowel portion is selected. The subject also reads a phonetically balanced text aloud at comfortable pitch and loudness. For AVQI version 03.01 in Dutch, the continuous-speech part is trimmed to the first 34 syllables of the text, the 3-second [a:] segment is appended, and the result is saved as a single sound file (Barsties and Maryn 2016). Acoustic analysis is applied only to the voiced segments of the continuous speech — extracted by an automated Praat detection script — plus the appended vowel segment (Barsties and Maryn 2016).
6.2.2 Component measures
All six parameters are obtained in Praat from the concatenated signal (Barsties and Maryn 2016; Delgado Hernández et al. 2018):
| Parameter | Definition |
|---|---|
| Smoothed cepstral peak prominence (CPPS) | Distance between the first rahmonic’s peak and the point of equal quefrency on the regression line through the smoothed cepstrum — see the CPPS chapter |
| Harmonics-to-noise ratio (HNR) | Base-10 logarithm of the ratio between periodic energy and noise energy, × 10 |
| Shimmer local (Shim) | Absolute mean difference between amplitudes of successive periods, divided by the average amplitude |
| Shimmer local dB (ShdB) | Base-10 logarithm of the difference between amplitudes of successive periods, × 20 |
| General slope of the spectrum (Slope) | Difference between the energy in 0–1000 Hz and in 1000–10000 Hz of the long-term average spectrum |
| Tilt of the spectral trendline (Tilt) | The same difference computed on the trendline through the long-term average spectrum |
6.2.3 Regression equation (version 03.01)
As printed by Barsties and Maryn (2016):
\[ \text{AVQI}_{03.01} = \bigl(4.152 - 0.177\,\text{CPPS} - 0.006\,\text{HNR} - 0.037\,\text{Shim} + 0.941\,\text{ShdB} + 0.01\,\text{Slope} + 0.093\,\text{Tilt}\bigr) \times 2.8902 \]
The published Praat script displays the resulting score on a 0–10 axis, marked green below the Dutch threshold of 2.43 and red above it (Barsties and Maryn 2016).
Two validation papers print this equation differently. Barsties v. Latoszek et al. (2020) print the HNR term as “(0.0006-HNR)” and Delgado Hernández et al. (2018) as “(0.0006 − HNR)”, both with a minus sign where a multiplication is expected; Delgado Hernández et al. (2018) also print the Tilt coefficient as 0.0993. The other sources — Barsties and Maryn (2016) (text and appendix script), Englert et al. (2021), and Castillo-Allendes et al. (2023) — give 0.006 × HNR and 0.093 × Tilt, as above. All printed forms were checked against the publisher PDFs, so the differences are in the publications. When implementing the index, take the coefficients from the Praat script printed in the appendix of Barsties and Maryn (2016) rather than from an equation retyped in a later paper.
6.2.4 Versions 02.02 and 03.01
In the earlier AVQI versions the analyzed continuous-speech part was much shorter than the constant 3-second vowel part (1.53 s of voiced speech in Dutch, 1.99 s in Japanese) (Delgado Hernández et al. 2018). Version 03.01 balances the analyzed durations of the two speech types by expanding the continuous-speech sample from 17–22 syllables to around 34 syllables, for higher ecological validity and balanced internal consistency (Barsties and Maryn 2016; Delgado Hernández et al. 2018). Early AVQI workflows required both SpeechTool and Praat; once CPPS was implemented in Praat, the index ran in Praat alone, with highly comparable outcomes (Barsties and Maryn 2016).
The regression-based combination of measures over a concatenated sample was explored for tracheoesophageal (alaryngeal) voice by Maryn et al. (2009), where a two-factor model combining CPP with the prominence of the second spectral harmonic reached rs = 0.87 with listener judgments.
6.3 What it captures
The AVQI quantifies overall voice quality as judged by the grade (G) of the GRBAS scale — the degree of hoarseness or voice abnormality (Barsties and Maryn 2016; Delgado Hernández et al. 2018). Concurrent validity with mean perceptual G is strong: r = 0.815 in the Dutch external validation (66.4% of variance) (Barsties and Maryn 2016) and r = 0.835 in Spanish (69.7% of variance) (Delgado Hernández et al. 2018). CPPS is the main factor in the multivariate model (Barsties and Maryn 2016). Barsties v. Latoszek et al. (2018) report, from the earlier overall-voice-quality meta-analysis of Maryn, Roy et al. (2009), weighted correlations for CPPS of 0.63 for sustained vowels and 0.88 for continuous speech.
A previous AVQI investigation (Maryn, De Bodt, et al. 2010) reported high sensitivity to voice changes through voice therapy (Barsties and Maryn 2016).
6.4 Normative data
Diagnostic thresholds separate normophonic from hoarse voices, with hoarseness defined as mean perceptual grade G ≥ 0.5 (Barsties and Maryn 2016). The threshold differs between AVQI versions and languages (Barsties v. Latoszek et al. 2020), so each row states its language and pipeline. Use the row validated for the patient’s language and the AVQI version in use; see Pipeline Dependence and Reference Values.
| Language | Sample | Pipeline | Threshold | Sensitivity | Specificity | AUC | Class |
|---|---|---|---|---|---|---|---|
| Dutch, external validation | 970 dysphonic, 88 normal | AVQI 03.01; Praat 5.3.57 | 2.43 | 0.785 | 0.932 | 0.923 | language-conditional |
| Spanish | 136 dysphonic, 47 normal | AVQI 03.01; Praat 6.0.22 | 2.28 | 0.748 | 0.946 | 0.905 | language-conditional |
| German | 175 dysphonic, 43 healthy; 27 syllables | AVQI 03.01; Praat 5.3.57 | 1.85 | 0.72 | 0.90 | 0.90 | language-conditional |
| French | 90 ENT patients, 30 controls; 27 syllables | AVQI 03.01; Praat version not stated; recorded at 11,025 Hz | 2.33 | 0.598 | 1.00 | 0.788 | language-conditional |
| Brazilian Portuguese | 113 dysphonic, 37 nondysphonic; counting 1–11 (17 syllables) | AVQI 03.01 script; Praat 6.0.40 | 1.33 | 0.788 | 0.906 | 0.904 | language-conditional |
| Brazilian Portuguese, counting 1–11 | 40 dysphonic, 13 healthy; 17 syllables | AVQI 03.01; Praat 6.0.06 | 1.16 | 0.86 | 0.80 | 0.87 | language-conditional |
| Brazilian Portuguese, reading text | same 53 speakers; 32 syllables | AVQI 03.01; Praat 6.0.06 | 1.56 | 0.886 | 1.00 | 0.963 | language-conditional |
| US English | 148 with voice disorders, 49 non-clinic-seeking; first 22 syllables of the Rainbow Passage | AVQI 03.01 in VOXplot 2.0.0; Praat 6.3.06 for extraction | 1.17 | 0.62 | 0.95 | 0.84 | language-conditional |
| Pooled: eight AVQI v03 studies (context only) | 12 languages | multiple | none; study thresholds ranged 1.33–3.15 | 0.82 | 0.92 | 0.94 (SROC) | meta-analytic |
The pooled row is descriptive context, not a cutoff. Its sensitivity was highly heterogeneous across studies (I² = 95.3%), its specificity much less so (I² = 16.4%) (Jayakumar and Benoy 2024).
Two further thresholds appear in the literature only second-hand and are not used as cutoffs here, because their pipeline is not stated where they are reported: 2.43 with sensitivity 0.936 and specificity 1.00 for an internal Dutch validation, as reported by Barsties and Maryn (2016), and 2.06 with sensitivity 0.721 and specificity 0.938 for Japanese, as reported by Delgado Hernández et al. (2018).
Likelihood ratios at the optimal threshold: LR+ = 11.54, LR− = 0.23 (Dutch external validation) (Barsties and Maryn 2016); LR+ = 13.85, LR− = 0.26 (Spanish, threshold chosen by the Youden index) (Delgado Hernández et al. 2018); LR+ = 7.40, LR− = 0.31 (German) (Barsties v. Latoszek et al. 2020); LR+ = 8.38, LR− = 0.23 (Brazilian Portuguese) (Englert et al. 2021); LR+ = 12.46, LR− = 0.40 (US English) (Castillo-Allendes et al. 2023). The LR+ values printed by Englert et al. (2020) are below 1 or negative and cannot be correct, so they are not reproduced here.
The original AVQI, which carried no version number, reported two cutoff options: 2.36 (sensitivity 91%, specificity 59%) and 2.95 (sensitivity 74%, specificity 96%). The authors suggest 2.95 may be more appropriate when the main interest is screening (Maryn, Corthals, et al. 2010). These values apply to that version only.
The US English validation standardized the continuous-speech part at 22 syllables of the Rainbow Passage, recorded in a quiet environment (ambient noise below 35 dB) with a head-mounted microphone at 44.1 kHz / 16 bit, and computed the index in VOXplot 2.0.0, which the authors state uses established Praat algorithms and the same equation (Castillo-Allendes et al. 2023). Its threshold, 1.17, lies below the range of the AVQI meta-analysis, whose literature search ended in July 2021 (Jayakumar and Benoy 2024). The authors note that other English dialects and accents remain to be examined (Castillo-Allendes et al. 2023).
6.5 Confounds and cautions
Language-specific standardization is required. The continuous-speech part must be trimmed to a syllable count that yields about 3 seconds of voiced speech, and the count differs by language: 34 syllables in Dutch, 33 in Spanish, 30 in Japanese for AVQIv3 (Delgado Hernández et al. 2018).
Thresholds differ by language. Reported optimal cutoffs are 2.43 (Dutch), 2.28 (Spanish), and 2.06 (Japanese). The differences are small, which may indicate that AVQI is relatively insulated from inter-language phonetic differences (Delgado Hernández et al. 2018).
Recording conditions matter. Validation recordings were made in sound- treated rooms with head-mounted condenser microphones at 44.1 kHz / 16 bit, with environmental noise verified against a signal-to-noise norm (Barsties and Maryn 2016; Delgado Hernández et al. 2018). The AVQI script itself states that reliable analysis requires optimal data-acquisition conditions (Barsties and Maryn 2016).
The Praat version matters. With the same AVQI 03.01 script and the same 218 German recordings, seven Praat versions formed three clusters. Praat 6.0.44 introduced a bug in the CPPS computation that 6.0.47 fixed. With version 6.0.46, AVQI scores were on average about 4.2 points lower, and sensitivity at the German threshold of 1.85 fell from 72% to 23%. Across the unaffected versions, mean differences were negligible, but single recordings differed by up to about 0.6 points (Stappenbeck et al. 2020). The authors recommend using the Praat version of the validation study for the language in question, or checking a new version against an earlier one on test recordings (Stappenbeck et al. 2020).
Sampling rate. The French validation recorded at 11,025 Hz (Pommée et al. 2020), so its signals contain nothing above about 5.5 kHz. The AVQI script computes the LTAS slope and trend-line tilt over bands that reach 10 kHz (Barsties and Maryn 2016), so two of the six components are computed on a narrower band than in the other validations. Treat the French threshold as applying to recordings made at that rate.
Report the version. Versions 02.02 and 03.01 differ in the length of the continuous-speech part (17 vs. 34 syllables) and in the regression weighting (Barsties and Maryn 2016). Report the version with every score.
Base rates affect ROC statistics. The Dutch external validation sample contained 8% normophonia and 92% dysphonia; likelihood ratios were used because they are less vulnerable to base-rate differences than ROC statistics alone (Barsties and Maryn 2016).
Voicing detection is part of the pipeline. Voiced segments of continuous speech are extracted automatically by a Praat script before measurement (Barsties and Maryn 2016).
6.6 Validation
Across 1118 subjects in two studies, AVQI 03.01 showed a homogeneous weighted mean correlation of r = 0.821 with perceptual overall voice quality, against r = 0.790 for the initial AVQI model pooled over 507 subjects (Barsties and Maryn 2016). In the same pooled analysis, the Dysphonia Severity Index reached a heterogeneous weighted mean r = 0.524 and the Cepstral Spectral Index of Dysphonia r = 0.788 (sustained vowel) and 0.748 (continuous speech); AVQI 03.01 reached the highest concurrent validity among these multiparametric indices (Barsties and Maryn 2016). Language validations include Dutch (Barsties and Maryn 2016) and Spanish (Delgado Hernández et al. 2018), with Japanese figures reported in the latter.
6.7 Compute it in PhonaLab
PhonaLab is developed by the author of this reference; see Competing interests.
Upload a sustained vowel and a continuous speech sample to compute AVQI on your own recording, with the version and language protocol reported alongside the score.