What is it about?
Machine learning models are often compared across many datasets, which can make benchmarking expensive and time-consuming. In this work, we study whether a much smaller set of carefully selected datasets can produce nearly the same overall ranking of models as the full benchmark. We compare several dataset selection strategies across time series classification, recommender systems, and natural language processing. Our results show that good selection methods can substantially reduce the number of datasets needed in some domains, but their success depends strongly on how well the datasets can be represented by informative features. We also provide a statistical evaluation framework for measuring how reliably a reduced benchmark preserves the conclusions of the full one.
Featured Image
Photo by Luke Jones on Unsplash
Why is it important?
Large benchmarks can be expensive to run, while using only a few datasets can change the final model ranking. Our approach selects representative datasets without knowing the results of the full benchmark, using only information available about the datasets in advance. This makes it useful when full evaluation is too expensive to perform.
Perspectives
For me, the key question is whether we can choose a smaller benchmark before running all the models on all the datasets. Our results show that this is possible in some domains, but the quality of the dataset representation is crucial.
Rostislav Gusev
Read the Original
This page is a summary of: Benchmarking on Tasks That Matter: Dataset Selection for Preserving Model Rankings, August 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3770855.3817569.
You can read the full text:
Resources
Contributors
The following have contributed to this page







