AI reveals how the brain distinguishes a voice amid noise
Researchers have developed an artificial neural network that mimics the human ability to pick out a single voice in a noisy environment. This model helps to better understand the mechanisms of auditory attention and could contribute to the creation of more effective hearing implants.
Cursus
Training an artificial neural network to identify specific sound features has enabled researchers to uncover the biological mechanisms underlying the human ability to focus on a single conversation in a noisy environment.
Studying the "Cocktail Party" Phenomenon
For decades, neurobiologists have investigated the so-called "cocktail party problem"—the brain’s capacity to selectively pick out one voice among a multitude of background noises. It was known that the brain achieves this by amplifying the activity of neurons that respond to certain audio cues. However, until recently, there was no computational model confirming that this mechanism alone is sufficient for real-world performance.
Development of an Artificial Neural Network
A team of researchers from the Massachusetts Institute of Technology developed an artificial neural network that mimics the human ear’s ability to distinguish individual voices. In a study published in Nature Human Behavior, they demonstrated that the brain uses a strategy known as multiplicative feature amplification. This means that when listening for a target voice, the brain boosts neural signals associated with that voice’s unique characteristics, such as its pitch, while simultaneously reducing the volume of competing sounds.
Model Validation
To test this hypothesis, the artificial model was presented with a short audio cue featuring a specific voice, followed by a noisy mixture of overlapping voices. The model successfully isolated the target voice from the background, showing results comparable to human performance in various conditions. Additionally, the system reproduced typical human hearing errors, such as difficulty distinguishing between two voices with similar pitch.
Advantages of the New Model
Previous models lacked a key human capability: they could not focus on a specific object or sound and adjust their response based on the chosen target. This limitation reduced their effectiveness.
The Impact of Spatial Arrangement
The model also allowed researchers to quickly assess how the spatial arrangement of sound sources affects perception. The system predicted that distinguishing voices is significantly easier when speakers are positioned horizontally rather than vertically—a phenomenon later confirmed by experiments involving human participants.
Application Prospects
Researchers believe that this model could contribute to the development of more advanced cochlear implants, enabling people to concentrate more effectively in noisy environments.
