UTD Press Journals
Convergence in Science and Society

Handling Missing Data in CALL - A Data Quality-Driven Imputation Framework for Learner Analytics: Robustness Under Distribution Shift

Read & download PDF
Abstract

Robustness depends on whether conclusions remain stable when the data distribution, case mix, prevalence, or operating environment differs from the reported setting. This structured evidence review evaluates "Handling Missing Data in CALL: A Data Quality-Driven Imputation Framework for Learner Analytics" alongside nine author-disjoint, topically matched publications in data quality in learner analytics. It compares construct definitions, evaluation choices, operating assumptions, and reported limitations instead of treating bibliographic similarity as empirical equivalence. Viewed through robustness under distribution shift, the map separates claims supported by the available record from questions that still require full-text extraction, replication, or new experiments. The synthesis is interpretive rather than meta-analytic and therefore does not present a pooled effect estimate or a new causal result. The resulting agenda uses prespecified shift scenarios, subgroup analysis, calibration checks, and post-deployment monitoring to locate failure boundaries.

Keywords
data quality in learner analyticsrobustness under distribution shiftevidence synthesisreproducibilityresearch evaluation
References
  1. Cao, X., Tao, J., Liu, Z., Lyu, R., & Li, J. (2026). Handling Missing Data in CALL: A Data Quality-Driven Imputation Framework for Learner Analytics. Future-Adaptive Intelligence and Lifelong Systems, 1(1).
  2. EL MOUDDEN, T., & Lachgar, N. (2026). When Missing Data Matters: Imputation, TOPSIS, and the Misrepresentation of Morocco’s Education System. . https://doi.org/10.2139/ssrn.6643909 DOI
  3. KALKAN, Ö.-K., KARA, Y., & KELECİOĞLU, H. (2018). Evaluating Performance of Missing Data Imputation Methods in IRT Analyses. International Journal of Assessment Tools in Education, 5(3), 403-416. https://doi.org/10.21449/ijate.430720 DOI
  4. Anand, V., & Mamidi, V. (2020). Multiple Imputation of Missing Data in Marketing. 2020 International Conference on Data Analytics for Business and Industry: Way Towards a Sustainable Economy (ICDABI), 1-6. https://doi.org/10.1109/icdabi51230.2020.9325602 DOI
  5. Xu, L., & Qiu, A. (2022). Multiple Imputation by Chained Equations for Missing Data in UK Biobank. 2022 6th Annual International Conference on Data Science and Business Analytics (ICDSBA), 72-82. https://doi.org/10.1109/icdsba57203.2022.00026 DOI
  6. Yang, L., & Chiang, J.-A. (2020). Use Case and Performance Analyses for Missing Data Imputation Methods in Big Data Analytics. Proceedings of 2020 6th International Conference on Computing and Data Engineering, 107-111. https://doi.org/10.1145/3379247.3379270 DOI
  7. Wang, K., Luo, M., Deng, M., & Chen, H. (2022). Nested Random Forest: A Personalized Imputation Method for Missing Data. 2022 8th International Conference on Big Data and Information Analytics (BigDIA), 119-126. https://doi.org/10.1109/bigdia56350.2022.9874127 DOI
  8. Sebastian, A.-M., Peter, D., & Sebastian, R.-A. (2025). An optimal imputation algorithm for reducing bias and errors in missing data handling for AI models. Decision Analytics Journal, 16, 100627. https://doi.org/10.1016/j.dajour.2025.100627 DOI
  9. He, Y., Zhang, G., & Hsu, C.-H. (2021). Multiple Imputation Analysis for Nonignorable Missing Data. Multiple Imputation of Missing Data in Practice, 375-406. https://doi.org/10.1201/9780429156397-13 DOI
  10. Sivakani, R., Rahila, J., Sudha, P., Priscila, S.-S., Shynu, T., Minu, M.-S., & Pradeep, V. (2025). A Smart Review on Imputation Techniques for Handling Missing Data. Machine Learning, Predictive Analytics, and Optimization in Complex Systems, 41-62. https://doi.org/10.4018/979-8-3373-5203-9.ch003 DOI
Publication details
Journal
Convergence in Science and Society
Volume
1 (2026)
Article number
css20260068
License
CC BY 4.0