UTD Press Journals
Computing Architectures and Data Systems

Reproducible Benchmarking of Machine-Learning Systems

Read & download PDF
Abstract

This review examines reproducible benchmarking of machine-learning systems. The organizing question is which controls make results portable across hardware, software versions, datasets, and tuning budgets. Ten related scholarly sources are synthesized through a decision-centered framework spanning problem definition, mechanism, measurement, evaluation, implementation, and governance. The review does not invent experiments, pooled estimates, or unreported quantitative results. It instead evaluates the strength and transferability of the available evidence, with particular attention to presenting a narrowly optimized speedup without uncertainty or complete configuration. The resulting framework links technical or empirical performance to explicit use conditions and identifies tests that should precede wider adoption in AI systems and infrastructure evaluation.

Keywords
machine-learning systemsbenchmarkingreproducibilityperformancemeasurement
References
  1. Azodi, C. B., Bolger, E., McCarren, A., Roantree, M., Campos, G. D. L., & Shiu, S. H. (2019). Benchmarking Parametric and Machine Learning Models for Genomic Prediction of Complex Traits. G3 Genes Genomes Genetics, 9(11), 3691-3702. https://doi.org/10.1534/g3.119.400498 DOI
  2. Borlido, P., Schmidt, J., Huran, A. W., Tran, F., Marques, M. A. L., & Botti, S. (2020). Exchange-correlation functionals for band gaps of solids: benchmark, reparametrization and machine learning. npj Computational Materials, 6(1). https://doi.org/10.1038/s41524-020-00360-0 DOI
  3. Duan, Y., Chen, X., Houthooft, R., Schulman, J., & Abbeel, P. (2016). Benchmarking Deep Reinforcement Learning for Continuous Control. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1604.06778 DOI
  4. Fu, X., Wu, Z., Wang, W., Xie, T., Keten, S., Gomez-Bombarelli, R., & Jaakkola, T. (2022). Forces are not Enough: Benchmark and Critical Evaluation for Machine Learning Force Fields with Molecular Simulations. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2210.07237 DOI
  5. Hannousse, A., & Yahiouche, S. (2021). Towards benchmark datasets for machine learning based website phishing detection: An experimental study. Engineering Applications of Artificial Intelligence, 104, 104347. https://doi.org/10.1016/j.engappai.2021.104347 DOI
  6. Hu, W., Fey, M., Zitnik, M., Dong, Y., Ren, H., Liu, B., Catasta, M., & Leskovec, J. (2020). Open Graph Benchmark: Datasets for Machine Learning on Graphs. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2005.00687 DOI
  7. Purushotham, S., Meng, C., Che, Z., & Liu, Y. (2018). Benchmarking deep learning models on large healthcare datasets. Journal of Biomedical Informatics, 83, 112-134. https://doi.org/10.1016/j.jbi.2018.04.007 DOI
  8. Wu, Z., Ramsundar, B., Feinberg, E. N., Gomes, J., Geniesse, C., Pappu, A. S., Leswing, K., & Pande, V. (2017). MoleculeNet: a benchmark for molecular machine learning. Chemical Science, 9(2), 513-530. https://doi.org/10.1039/c7sc02664a DOI
  9. Wójcikowski, M., Ballester, P. J., & Siedlecki, P. (2017). Performance of machine-learning scoring functions in structure-based virtual screening. Scientific Reports, 7(1), 46710. https://doi.org/10.1038/srep46710 DOI
  10. Zech, J. R., Badgeley, M. A., Liu, M., Costa, A. B., Titano, J. J., & Oermann, E. K. (2018). Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study. PLoS Medicine, 15(11), e1002683. https://doi.org/10.1371/journal.pmed.1002683 DOI
Publication details
Journal
Computing Architectures and Data Systems
Volume
1 (2026)
Article number
cads20260004
License
CC BY 4.0