Hearing Mathematics · How Sound Enters the Body and Mind · Article Nine
On the page, a four-voice fugue appears clearly divided. A subject enters in one voice, another answers, and later material crosses among the parts. By the time sound reaches an ear, there are no four independent channels. Pressure variations have already added together, and the eardrum receives one waveform changing through time. A listener can nevertheless select one voice, follow it through interference from other notes, and then move attention to another line.
The capacity is not peculiar to music. Following one speaker in a restaurant, or separating vehicles, footsteps and birds in a street, requires an auditory system to infer sources from mixtures. Music turns this auditory scene analysis into an art that can be deliberately designed. Voices must differ enough to remain separable and relate enough to form one work.
Once waves are added, the physical world does not label their sources
When two instruments play, sound pressure at a point is approximately the sum of their waveforms. A microphone records the sum but does not automatically mark which portion belongs to violin and which to oboe. A Fourier transform can reveal frequency components without determining from a single spectrum how many sources produced them. Harmonics from different sources can overlap, while components of one source change through time.
Hearing uses cues to construct the most plausible grouping. Components that begin and end together tend to belong to one object. Frequencies moving smoothly are more readily connected. Events close in pitch, similar in timbre and continuous in apparent location tend to form a stream. Large pitch separation, high rate or sudden timbral change can promote separation.
The classic repeated ABA sequence makes the process audible. When A and B are close in pitch and proceed slowly, the sequence often forms one galloping pattern. Increase frequency separation or speed and the A tones may form one stream while the Bs form another. The physical sequence has acquired no new source, but its perceptual organisation has changed. A review of computational auditory scene analysis describes an intermediate region in which one stimulus can alternate between integrated and segregated percepts. Attention can bias the result without exercising unlimited control.
A musical “voice” is not one pre-existing line in the spectrum
A voice in music theory is a directed and relatively independent musical line. It need not correspond to one singer or instrument. A piano can perform four voices in one timbre. A solo violin can imply several lines through double stops and compound melody. Several orchestral instruments can combine into one functional voice.
Voice identity therefore arises through continuity and structural role, not only source identity. A fugue subject stated first in an inner voice and later in the soprano remains recognisably the same subject type, but not one continuously sounding source. Auditory organisation and musical memory overlap: one links adjacent events into a stream; the other recognises shapes repeated across longer time.
Pitch proximity normally supports continuity. A large leap risks having a note captured by a neighbouring line. Voices moving in the same direction at the same time can fuse; independent rhythms, staggered onsets and separated registers help them remain distinct. The practices of voice leading are not merely rules visible on paper. Many are related to whether lines can be perceptually followed.
Bottom-up cues and attention work together
When flute and cello occupy distinct registers, both timbre and pitch range support separation. These are comparatively bottom-up cues. When two piano voices share a register and move in synchrony, the boundary is harder to preserve. A listener can use prior knowledge of a subject, a score or deliberate attention to one range to improve tracking.
Experiments using polyphonic music to investigate auditory scene analysis emphasise that segregation and integration are not mutually exclusive tasks. A listener must preserve horizontal continuity within each line while hearing vertical harmonic relations among them. Segregation alone would yield unrelated sounds; integration alone would erase polyphonic independence.
Attention cannot arbitrarily rewrite acoustics. When components are perfectly synchronous, co-located and highly similar, willpower may not split them into sources. When frequency separation is very large and presentation extremely fast, it can be difficult to force a single smooth melodic hearing. Perception is not free interpretation. It selects a sustainable organisation among physical cues, learned models and present tasks.
Bach’s The Art of Fugue: how one timbre can still support many lines
Bach’s The Art of Fugue develops a central subject through a sequence of contrapuntal works. It does not specify one comprehensive instrumentation, and keyboard, string quartet and other realisations reveal its organisation differently. The scores and work resources show the subject in original and inverted forms, augmentation and increasingly complex combinations.
Begin with an earlier Contrapunctus and tap lightly at each subject entry. On the first hearing, recognise the subject’s contour without trying to follow a complete part. On the second, choose the highest voice and connect each nearby note, even when the subject appears elsewhere. On the third, follow the bass. Changes of attention give one recording different foregrounds, while unattended voices continue to shape harmony and density.
A string-quartet performance supplies timbral and spatial differences that help separate voices. A keyboard performance reduces those cues, making register, articulation, dynamics and phrasing more important. One realisation is not automatically more correct and the other more pure. Instrumentation changes the evidence available to auditory scene analysis. The compositional structure persists while the perceptual route into it changes.
An experiment that used excerpts from The Art of Fugue to study tracking of musical voices found that timbral heterogeneity affected performance differently across listener groups, with age and hearing-aid status altering access to cues. The result is a useful warning. Four objectively notated voices do not guarantee that every listener in every instrumentation hears four equally clear streams.
Difficulty hearing a voice is not always failure to pay attention
Hearing loss can reduce frequency selectivity and temporal detail, making nearby components harder to separate. Reverberation smears onsets and offsets. Collapsing a stereo recording to mono removes spatial cues. Performance rate, voice balance and room acoustics all affect the audibility of polyphony.
It is unfair, then, to attribute difficulty following complex music simply to poor cultivation. Understanding depends on perceptual evidence being available. Performers use phrasing, articulation, level and tempo to expose structure. Recording engineers and venues preserve or mask scene cues. Hearing technologies also face difficult compromises between speech intelligibility and musical spectra.
Familiarity changes what can be heard. On first exposure, a fugue subject may be clear only at its opening. After reading the score or singing one part, entries previously absorbed into the texture can suddenly become apparent. The sound file has not changed. The internal model available for attention has.
Polyphony exists among score, sound wave and perceptual organisation
The statement “there are four voices here” may refer to three different facts: the composition contains four lines; the performance produces four related streams of action; and the listener’s experience contains four trackable objects. They normally correspond, but not automatically. A score can be clearly divided while poor balance leaves a block of chords. A monophonic instrument can imply two virtual voices through register leaps and timing.
Mathematics can represent frequencies, onsets, intervals and correlations, and algorithms can attempt source separation. Every separation still embodies assumptions. Do partials sharing one fundamental count as one source? Is a subject transferred to another instrument still the same object? A musical voice cannot be defined by signal statistics alone because it has syntactic and compositional roles.
The best conclusion is neither that ears faithfully copy a score nor that voices are merely subjective inventions. A voice is a stable object formed jointly by physical cues, composition, performance and a listener’s attention. It is objective enough to be tracked and analysed by many people, yet conditional enough for clarity to change with instrumentation, hearing and learning.
While listening to The Art of Fugue, allow yourself to lose a line and find it again. The value of polyphony is not only the simultaneous possession of all information. Foreground and background can exchange roles. Many sounds reach one eardrum; hearing does not permanently separate them. It establishes a movable order in which one path can be chosen while the others continue.
Primary sources and further listening
- Snyder and Alain, “Computational Models of Auditory Scene Analysis: A Review”
- Research on stream segregation and integration using polyphonic music
- Research on tracking musical voices in Bach’s The Art of Fugue
- Bach, The Art of Fugue: scores and recordings
Continue reading: Explore the Hearing Mathematics series.
If you would like to bring these ideas about listening, understanding, and practice to the keyboard, you might try ScoreFlow, an app I developed to make score reading and daily practice flow more naturally together.
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.