IISc research paper on AI training efficiency among top 15 at global CVPR conference

PC : linkedin/Indian Institute of Science (IISc)
Bengaluru: A research paper by scientists from the Indian Institute of Science (IISc), Bengaluru, highlighting a new approach that could significantly reduce the cost and environmental impact of training artificial intelligence (AI) models, has secured a place among the top 15 submissions at the Computer Vision and Pattern Recognition (CVPR) 2026 conference held in Colorado, US.
The paper, titled 'Rethinking Dataset Distillation: Hard Truths about Soft Labels', was authored by Priyam Dey, R Venkatesh Babu, Additya Sahdev, Sunny Bhati and Konda Reddy Mopuri from the IISc Department of Computational and Data Science (CDS).
According to the IISc, the paper reached the finals at the conference after being shortlisted from nearly 16,000 submissions received for the annual event, one of the world's leading forums for research in computer vision and pattern recognition.
Explaining the work, CDS Head Prof. R Venkatesh Babu was quoted by The Indian Express as saying that the study revisits the concept of 'dataset distillation', which seeks to reduce the amount of data required to train AI systems without compromising their performance.
He said, "With AI, we have a large amount of data used in training models… you may need a very large and expensive network of training data. In a dataset, where so many samples are available, can we get a handful of samples with which we can train the AI model? Then the training cost can come down drastically."
The research challenges prevailing assumptions in AI model training by suggesting that carefully selected random samples can perform on par with more complex dataset distillation techniques.
"This is kind of a deviation from what people were doing continuously. We wanted to look back and say, this is not the correct way. Random (training data) samples also give you the same accuracy," Prof. Babu said.
He noted that the current work has been demonstrated on image classification tasks, such as categorizing nearly a million images into around 1,000 classes, and pointed out that the underlying approach could also be extended to other forms of data, including audio.
Prof. Babu further said the findings could also contribute to making AI systems more sustainable by lowering the computational resources required for training. "The volume of data is so much that we never pay attention to it and feed whatever is available to the machinery… that is what is emitting huge amounts of carbon. Any effort to reduce this amount of data could significantly reduce the carbon footprint," he said.
He added that reducing the amount of training data would also lower infrastructure requirements, including electricity consumption, graphics processing units (GPUs) and other computing resources, making AI development more cost-effective.



