UTD Press Journals
Computing Architectures and Data Systems

From Code to Correctness - Closing the Last Mile of Code Generation with Hierarchical Debugging: Evaluation Design and Construct Validity

Read & download PDF
Abstract

Evaluation is persuasive only when the measured outcome corresponds to the construct claimed by the study and the comparison answers the stated research question. This structured evidence review evaluates "From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging" alongside nine author-disjoint, topically matched publications in AI-assisted software engineering. It compares construct definitions, evaluation choices, operating assumptions, and reported limitations instead of treating bibliographic similarity as empirical equivalence. Viewed through evaluation design and construct validity, the map separates claims supported by the available record from questions that still require full-text extraction, replication, or new experiments. The synthesis is interpretive rather than meta-analytic and therefore does not present a pooled effect estimate or a new causal result. The resulting agenda aligns claims, outcomes, comparators, sampling, and uncertainty before any performance estimate is interpreted.

Keywords
AI-assisted software engineeringevaluation design and construct validityevidence synthesisreproducibilityresearch evaluation
References
  1. Shi, Y., Wang, S., Wan, C., Wang, M., & Gu, X. (2026). From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging. In 2026 IEEE/ACM International Conference on Software Engineering (ICSE).
  2. Nguyen, X.-T. (2019). Taxing Facebook Code: Debugging the Tax Code and Software. . https://doi.org/10.31228/osf.io/kanc4 DOI
  3. V. Saravanan,, S. Kavitha,, S. Ravi,, A. Seetha,, Ch Rambabu,, & Tatiraju V. Rajani Kanth, (2025). Generative AI in Software Engineering: Revolutionizing Code Generation and Debugging. International Journal of Computational and Experimental Science and Engineering, 11(2). https://doi.org/10.22399/ijcesen.1718 DOI
  4. Vikram, M., Eluri, N., Dundi, U., Velicharla, R., Surapuraju, S., & Kondapureddy, V.-R. (2026). Advanced artificial intelligence algorithms for software engineering automating code generation, debugging, and software maintenance. AIP Conference Proceedings, 3418, 050039. https://doi.org/10.1063/5.0342075 DOI
  5. Gülmez, B. (2026). Code generation with large language models: a survey from neural program synthesis to autonomous software development. Applied Intelligence, 56(6). https://doi.org/10.1007/s10489-026-07230-0 DOI
  6. Adnan, M., NOSCHANG KUHN, C.-C., & Xu, Z. (2025). Large Language Model Guided Self-Debugging Code Generation. . https://doi.org/10.2139/ssrn.5396508 DOI
  7. Li, S., Xie, K., Li, Y., Li, H., Ren, Y., Sun, L., & Zhu, H. (2025). TransferFuzz-Pro: Large Language Model Driven Code Debugging Technology for Verifying Propagated Vulnerability. IEEE Transactions on Software Engineering, 51(8), 2396-2411. https://doi.org/10.1109/tse.2025.3584774 DOI
  8. Lin, F., Kim, D.-J., & Chen, T.-H. (2025). SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents. 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), 1527-1539. https://doi.org/10.1109/icse55347.2025.00140 DOI
  9. Wang, F., Xi, X., Cui, Z., Dai, H., & Wang, X. (2025). Embedding Traceability in Large Language Model Code Generation: Towards Trustworthy AI-Augmented Software Engineering. Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, 1760-1763. https://doi.org/10.1145/3696630.3730569 DOI
  10. Rose, L. (2020). An Efficient Transformer-Based Model for Automated Code Generation: Leveraging Large Language Models for Software Engineering. International Journal of Emerging Research in Engineering and Technology, 1, 1-9. https://doi.org/10.63282/3050-922x.ijeret-v1i3p101 DOI
Publication details
Journal
Computing Architectures and Data Systems
Volume
1 (2026)
Article number
cads20260006
License
CC BY 4.0