A silicone collar transforms silent speech into voice
South Korean researchers have developed a silicone collar that detects neck movements when words are mouthed silently and, using AI, converts them into speech that closely resembles the user's own voice. This technology could be used in medicine and in extremely noisy environments where traditional microphones are ineffective.
Ingenium
Researchers from POSTECH University of Science and Technology in South Korea have developed a silicone collar capable of detecting the slightest neck movements during silent speech and converting them into audible speech, reproducing the user's voice for the listener.
How the Device Works
The device is based on the principle that forming words is accompanied by distinctive movements of the neck muscles and skin, which create unique patterns for each syllable. Previously, capturing these signals relied on electromyography (EMG) and electroencephalography (EEG), but such methods required bulky equipment and were inconvenient outside laboratory settings.
The POSTECH team took a different approach: their collar combines soft silicone, a miniature camera, and motion sensors, all working together with artificial intelligence trained on the user's voice. The multi-directional sensor tracks the degree and direction of skin deformation during articulation, providing a detailed picture of mouth and throat activity. Markers applied to the silicone collar allow the camera to measure these deformations in real time.
The algorithm automatically adjusts for minor differences in the device's position each time it is worn, ensuring consistent readings. The detected deformation patterns are sent to an AI model, which determines the spoken word.
Testing and Effectiveness
During trials, the system was trained on words from the NATO phonetic alphabet, specifically designed for clarity in challenging conditions. When working with 26 words, the recognition accuracy reached 85.8%. After a word is recognized, it is transmitted wirelessly to a server, where a personalized text-to-speech model synthesizes an audio file that closely matches the user's real voice. Training the voice model requires less than 10 minutes of recordings.
The system demonstrated strong resistance to interference: even at noise levels around 90 dB (comparable to a construction site), the signal-to-noise ratio remained up to 33.75 dB, surpassing commercial EMG systems under similar conditions.
Limitations and Future Prospects
The current version of the device works only with a fixed vocabulary of 26 predefined words; free speech is not yet possible. Accuracy can drop to 39.72% if the user moves or turns their head actively. Future plans include expanding the vocabulary, testing with more users, and improving compensation for body movements.
Potential Applications
This technology could be used not only in medicine, for example, to help patients after laryngectomy, but also in environments where conventional microphones are ineffective or unavailable: industrial sites, emergency services, aviation, maritime operations, and military scenarios. The system was tested not only in white noise but also during demonstrations with a gas gun, where both noise and vibration were present.
Comparison with Other Developments
Similar approaches have previously been tested in laboratories, such as at Cambridge University, where collars with sensors were also used to detect throat vibrations during silent speech. The Cambridge prototype achieved a speech decoding accuracy of 95.25% and was not limited to a specific vocabulary. In recent studies, their system could also determine the user's emotional state. The unique feature of the POSTECH development is the use of artificial intelligence to recreate the user's individual voice.
