Yandex has introduced a method to improve voice commands.
Yandex researchers have proposed a new modular method that allows voice commands to be added to smart devices without compromising the recognition of already known phrases. This solution reduces the risk of "catastrophic forgetting" and is suitable for devices with limited resources.
Crius
Researchers at Yandex have developed a new method to address the issue of "catastrophic forgetting" when updating trigger words for smart devices. This approach enables the recognition of new voice commands without losing the ability to identify previously known phrases. The method will be presented at the international Interspeech 2026 conference, which will take place in Sydney from September 27 to October 1.
The Problem of Catastrophic Forgetting
Manufacturers of smart speakers, in-car assistants, and other voice-activated helpers regularly expand the list of commands that devices can recognize autonomously, without an internet connection. To achieve this, models are typically retrained on new data. However, after such updates, devices sometimes become less accurate at recognizing commands they previously understood. In machine learning, this phenomenon is known as "catastrophic forgetting." To mitigate this, methods of continual learning, such as elastic weight consolidation (EWC), are used, but these only reduce the problem rather than eliminate it entirely.
A New Modular Approach
The method developed at Yandex is based on modular model expansion. Instead of modifying the entire neural network, a small additional trainable module is added, responsible for recognizing new commands. This module uses features extracted by the main model and has its own classifier for the new commands. Meanwhile, the parameters of the base model, including normalization and the main classifier, remain unchanged. This approach preserves the accuracy of recognizing existing commands, while new commands are added by training only the additional module, without reconfiguring the entire system. The increase in neural network size is less than 10%.
Comparison with Other Methods
Experiments on the open Google Speech Commands dataset demonstrated the effectiveness of the new approach. The model was trained on new commands, and recognition of previously known words did not deteriorate. The new method reduced the average false reject rate (FRR) for new commands from 6.46% to 4.37% compared to a separate model of similar size. It also outperformed adapter and LoRA methods. With full retraining of the entire model, new commands were learned well, but the accuracy for old commands dropped significantly: the average rate of forgotten old commands increased from 2.71% to 69.08%.
Applications
This approach can be used in smart speakers, household appliances, in-car voice assistants, and other devices with limited computing resources.
