ICML 2026: New Methods Accelerate AI Training
At the ICML conference in Seoul, new methods were presented that accelerate AI training, improve the efficiency of computational resource usage, and enable work with limited or complex data. The research covers model optimization, faster search and recommendation systems, as well as reducing reliance on manual data labeling.
Cursus
Today, the artificial intelligence industry is focused on improving the efficiency of computational resource usage, accelerating model training, handling complex data, and reducing reliance on costly manual data labeling. These topics are the subject of a series of scientific papers presented by researchers and engineers at the International Conference on Machine Learning (ICML), held in Seoul, South Korea, from July 6 to 11.
International Conference on Machine Learning
ICML is considered one of the largest global conferences in the field of machine learning, annually bringing together experts from leading universities, research centers, and technology companies. In 2026, a total of 23,918 scientific papers were submitted for review, with 6,352 articles (26.6% of all submissions) accepted into the main program.
Efficient Use of Computational Resources
Modern AI developers strive to boost model performance without increasing computational power. One of the featured papers introduced new software modules for graph neural networks that analyze not only objects but also the relationships between them (for example, between users and products, documents, or road network segments). These modules enable much more efficient use of GPU memory: experiments showed up to 8.5 times faster computations, peak memory usage reduced by up to 76 times, and certain operations sped up by 3.9–10 times. The module code has been released as open source. The work received Spotlight status, awarded to papers with high program committee ratings (in 2026, 536 papers—2.2% of all submissions—received this status).
Accelerating Large Language Model Training
Another study addressed the challenge of speeding up large language model (LLM) training using pipeline parallelism. In this approach, some GPUs remain idle while waiting for others. Although asynchronous schemes can eliminate idle time, it was previously believed that gradient delay made them unstable for LLMs. Research showed that quality degradation is not due to delay, but to the choice of optimization algorithm. Modern methods like Muon handle gradient delay better than the classic AdamW. Additionally, a lightweight correction at the optimizer step level was implemented, allowing the asynchronous method to achieve quality comparable to synchronous training on Mixture of Experts (MoE) models with 10 billion parameters and training on 200 billion tokens.
New Optimization Algorithms
Another paper proposed two new optimization algorithms—SoftSignum and SoftMuon—which determine how a model updates its parameters during training. In experiments, these methods consistently outperformed several popular approaches, including AdamW.
Working with Complex Data
For graph-based tasks, the GraphPFN model was developed and pre-trained on over 1.6 million synthetic graphs. It demonstrated high performance even without additional fine-tuning, and when adapted to specific tasks, outperformed the approaches considered in the study on most real-world datasets. This method enables faster creation of models for new tasks and requires less data for training.
Another study showed that modern neural networks work more effectively with tabular data when they account for uncertainty in the data. To this end, a more efficient way of representing numerical features was proposed.
Training Models with Limited Data
In many applied tasks, there is enough data but a lack of high-quality labeling, especially in fields requiring expert input (such as medicine and industry). To address this, a method was proposed that allows the use of a small amount of labeled data in combination with a large volume of unlabeled data. This helps train models when labeling new data is too costly or time-consuming.
Accelerating Search and Recommendations
Search and recommendation systems process large volumes of data daily. A new method was proposed to speed up this work: in cases where the most accurate model is too resource-intensive to apply to all options, the new approach allows for pre-selecting the most suitable candidates, which are then evaluated by the more accurate model. This reduces computational costs.
