7  Harmonic-to-Noise Ratio (HNR)

Note

Living draft. This chapter reflects the sources cited below and will be revised as further sources are incorporated. Found an error? Use the Report an issue link in the sidebar.

7.1 Definition

A characteristic feature of hoarseness is the replacement of harmonics by noise energy: aperiodic sound intensifies at the expense of the periodic signal. It is therefore reasonable to conclude that the best index of hoarseness might be the ratio of one to the other, and that is what the harmonic-to-noise ratio measures (Baken and Orlikoff 2000).

The conceptual basis is simple. The voice is considered to have two components, perfectly periodic waves and random noise. The H/N ratio is the mean amplitude of the wave divided by the mean amplitude of the isolated noise components for the train of waves, and for convenience it is expressed in decibels (Baken and Orlikoff 2000).

7.2 Computation

The ratio approach was taken first by Kojima and colleagues and by Kitajima, but their computational procedure was complex and inconvenient. It was soon improved upon by Yumoto and colleagues, whose procedure was to objectify and quantify the features that appear in the spectrogram of the hoarse voice (Baken and Orlikoff 2000).

Neutralized noise energy (NNE) is a related measure with a different strategy. Regular changes over the sampling interval influence the assessment of noise, analogous to the long-term drift problems that plague jitter and shimmer. NNE avoids that contaminating influence by basing its analysis on relatively few vocal periods and detecting the noise component of the spectrum with a specially designed adaptive comb filter. Like the H/N ratio it is expressed in decibels (Baken and Orlikoff 2000).

An older family of measures quantifies spectral noise directly rather than as a ratio. In one such method, speakers sustained vowels for 7 s at a monitored 75 dB SPL; a 2-second segment was looped and repeatedly scanned by a spectrum analyzer at a very narrow 3 Hz bandwidth, and the minimum noise value was measured for each 100-Hz segment of the spectrum from 200 to 8000 Hz (Baken and Orlikoff 2000).

7.3 Normative data

This reference publishes no HNR cutoff.

HNR is strongly implementation-dependent, and no source reviewed here states a Praat version and settings for an HNR reference value. Under the rule in Pipeline Dependence and Reference Values, the absence is declared rather than filled.

Table 7.1: HNR reference values for a Praat-based pipeline.
Population Task Pipeline Cutoff Class
Any, measured in Praat Sustained /a/ — No published reference —

The values below come from other pipelines and are shown for context only.

Table 7.2: HNR and NNE values from other pipelines, for context only. Yumoto and spectral-noise rows: Baken and Orlikoff (2000). Age-group row: Ferrand (2002). Brazilian and NNE rows: Behlau (n.d.).
Population Task Value Pipeline Class
22 men and 20 women with no demonstrable vocal disorder Sustained /a/ Mean H/N 11.9 dB (SD 2.32), range 7.0–17.0; sexes did not differ significantly; 95 per cent confidence limit 7.4 dB Yumoto’s method; software and settings not stated unclassified
12 men and 8 women with various laryngeal pathologies, before surgery Sustained /a/ Mean H/N 1.6 dB; after surgery 11.3 dB (SD 3.13), range 5.9–17.6, with about 95 per cent reaching the expected normal range As above unclassified
14 women per group, three age groups Sustained vowel, onset and offset discarded Young 7.82 dB, middle-aged 7.86 dB, elderly 5.54 dB; the elderly group differed significantly from both younger groups Kay Elemetrics CSL 4300; Yumoto’s method except that the signal is not preemphasized; 51.2 kHz, 2-second frame; MiniDisc at 44.1 kHz unclassified for a Praat pipeline
Brazilian women and men, modal register Sustained vowel Women 9.4 dB; men 8.6 dB Soundscope 2.0; algorithm and settings not stated; reported second-hand unclassified
Women and men, modal register Sustained vowel Women 13.9 dB; men 11.8 dB Soundscope 2.0; second-hand unclassified
Women and men, falsetto register Sustained vowel Women 15.6 dB; men 15 dB Soundscope 2.0; second-hand unclassified
Pulse (basal) register Sustained vowel Not reliably computable: the noise component is high Soundscope 2.0; second-hand unclassified
NNE, normal limit Not stated Normal up to −10 dB; values nearer zero, such as −5 or −3 dB, strongly indicate phonatory aperiodicity Not stated unclassified

7.3.1 Where the 7 dB bound comes from

The widely repeated figure traces to Yumoto’s series, whose 95 per cent confidence limit for normal speakers was 7.4 dB, with a normal mean of 11.9 dB (Baken and Orlikoff 2000). The same 7.4 dB figure is reported independently (Ferrand 2002). It is commonly carried rounded to 7 dB, with the stronger claim that a ratio below 7 dB is necessarily pathological (Behlau, n.d.).

Three cautions follow. The bound belongs to the Yumoto algorithm family and is not a cross-system constant. In Yumoto’s own data, three of the twenty preoperative dysphonic subjects had ratios above the 7.4 dB limit, and they were the ones with very slight clinical hoarseness (Baken and Orlikoff 2000). And in the age-group series above, the healthy young and middle-aged means of 7.82 dB and 7.86 dB sit only just above the bound, while the elderly group’s mean of 5.54 dB falls below it although the group was normally speaking (Ferrand 2002). A 7 dB floor applied to older speakers would classify healthy voices as pathological.

7.4 What it captures

HNR rises with better harmonic structure and falls with more noise. It is lower in men and higher in women, and highest in falsetto, then modal, then pulse register — so a value is interpretable only in the register in which the voice was recorded (Behlau, n.d.).

Diffuse mass lesions can give paradoxically high values: in Reinke’s oedema the preserved harmonic component can keep the ratio high, so HNR alone can under-flag these voices. Small glottal gaps can give a low ratio or raised NNE, without exact correlation to the perceptual grade of dysphonia. The measure is useful for monitoring the course of paralytic dysphonias and partial laryngectomies (Behlau, n.d.).

NNE discriminates dysphonic voices better than HNR. In three groups of women — complete glottal coaptation, mid-posterior triangular gap, and vocal nodules — NNE was the most sensitive acoustic parameter and HNR did not differentiate the groups; F0 and jitter were also similar across groups, while shimmer was significantly higher in the nodule group (Behlau, n.d.). Published NNE data remain scarce, but preliminary reports suggest the measure performs well in discriminating abnormal voices (Baken and Orlikoff 2000).

What NNE tracks perceptually. In 26 girls and 24 boys aged ten, normal and dysphonic, NNE correlated most strongly with perceived breathiness, .78 in girls and .77 in boys; less with hoarseness, .50 and .57; and least with roughness, .38 and .48 (Baken and Orlikoff 2000).

The two measures are complementary: HNR appears more informative in normal voices, NNE in dysphonic ones (Behlau, n.d.).

7.5 Confounds and cautions

The most recording-sensitive of the common measures. In a three-system comparison, only fundamental frequency was resistant to differences in recording condition, and the harmonic-to-noise ratio was the most sensitive parameter of those examined (Behlau, n.d.).

System dependence. Extraction methods differ across programs and were often unstated in older studies, so values from different systems are not directly comparable; the software and version must be reported (Behlau, n.d.). The general warning in the English textbook is that different analysis systems may produce dissimilar results on the same signal, and that until inter-laboratory standards are agreed upon, comparison of data sets must be undertaken with caution (Baken and Orlikoff 2000). See Pipeline Dependence and Reference Values.

Spectral noise alone is a weak discriminator. Measured directly rather than as a ratio, spectral noise has been found to be only a fair differentiator of rough and abnormal voices (Baken and Orlikoff 2000).

Harmonic amplitude does not fall indefinitely as noise rises. As interharmonic spectral noise increases, energy at the harmonic frequencies does decrease, but only up to a point (Baken and Orlikoff 2000).

Vowel and fundamental frequency both matter. Vowels differ in median noise level in a consistent order, and among normal children a high fundamental frequency was associated with lower noise levels — though that relationship did not hold for the dysphonic children in the same study (Baken and Orlikoff 2000).

Signal typing applies. As with jitter and shimmer, HNR is reliably measurable only on Type 1 signals (Titze 1995).

7.6 Points of disagreement

How firm is the 7 dB floor. The Behlau chapter states that a ratio below 7 dB is necessarily pathological (Behlau, n.d.). Yumoto’s own series, the origin of the figure, found three preoperative dysphonic subjects above the 7.4 dB limit (Baken and Orlikoff 2000), and the normally speaking elderly group above averaged 5.54 dB, below it (Ferrand 2002). The bound is a 95 per cent confidence limit on one algorithm and one sample, not a diagnostic constant.

Cohort and software both move the normal range. The two sets of Brazilian modal-register means in Table 7.2 differ by roughly 3 to 4 dB despite both using the same program on Brazilian speakers (Behlau, n.d.). The sources reviewed here do not explain the gap. Any normative threshold used clinically should be tied to its extraction system and reference cohort rather than treated as universal.

Whether HNR or NNE should be the primary clinical noise measure remains open. The sources reviewed so far favour NNE for dysphonic voices and HNR for normal-voice screening (Behlau, n.d.).

7.7 Compute it in PhonaLab

Open PhonaLab →

PhonaLab is developed by the author of this reference; see Competing interests.

PhonaLab reports HNR with its algorithm, settings and task stated alongside the value. It deliberately shows no normal band and no severity band for HNR, because no published reference value is compatible with its pipeline — see Table 7.1. In particular it does not apply a 7 dB floor.

Baken, Ronald J., and Robert F. Orlikoff. 2000. Clinical Measurement of Speech and Voice. 2nd ed. Singular Thomson Learning.
Behlau, Mara. n.d. Voz: O Livro Do Especialista. Editora Revinter.
Ferrand, Carole T. 2002. “Harmonics-to-Noise Ratio: An Index of Vocal Aging.” Journal of Voice 16 (4): 480–87. https://doi.org/10.1016/S0892-1997(02)00123-6.
Titze, Ingo R. 1995. Workshop on Acoustic Voice Analysis: Summary Statement. National Center for Voice; Speech.