UTD Press Journals
Computing Architectures and Data Systems

Compiler Optimization for Heterogeneous Computing

Read & download PDF
Abstract

This review examines compiler optimization for heterogeneous computing. The organizing question is how compiler transformations should balance portability, performance, energy, and numerical fidelity across devices. Ten related scholarly sources are synthesized through a decision-centered framework spanning problem definition, mechanism, measurement, evaluation, implementation, and governance. The review does not invent experiments, pooled estimates, or unreported quantitative results. It instead evaluates the strength and transferability of the available evidence, with particular attention to reporting device-specific tuning as a portable compiler improvement. The resulting framework links technical or empirical performance to explicit use conditions and identifies tests that should precede wider adoption in scientific, data-intensive, and machine-learning workloads.

Keywords
compilersheterogeneous computingacceleratorsoptimizationportability
References
  1. Ashcraft, M. B., Lemon, A., Penry, D. A., & Snell, Q. (2017). Compiler Optimization of Accelerator Data Transfers. International Journal of Parallel Programming, 47(1), 39-58. https://doi.org/10.1007/s10766-017-0549-3 DOI
  2. Chang, A. X. M., Zaidy, A., Vitez, M., Burzawa, L., & Culurciello, E. (2019). Deep neural networks compiler for a trace-based accelerator. Journal of Systems Architecture, 102, 101659. https://doi.org/10.1016/j.sysarc.2019.101659 DOI
  3. Hayashi, A., Shirako, J., Tiotto, E., Ho, R., & Sarkar, V. (2016). Exploring compiler optimization opportunities for the OpenMP 4.x accelerator model on a POWER8+GPU platform. Scholarly publication, 68-78. https://doi.org/10.5555/3019120.3019127 DOI
  4. Liu, Z., Leng, J., Lu, G., Wang, C., Chen, Q., & Guo, M. (2020). Survey and design of paleozoic: a high-performance compiler tool chain for deep learning inference accelerator. CCF Transactions on High Performance Computing, 2(4), 332-347. https://doi.org/10.1007/s42514-020-00044-7 DOI
  5. Park, J., Yu, M., Kwon, J., Park, J., Lee, J., & Kwon, Y. (2024). NEST‐C: A deep learning compiler framework for heterogeneous computing systems with artificial intelligence accelerators. ETRI Journal, 46(5), 851-864. https://doi.org/10.4218/etrij.2024-0139 DOI
  6. Tiotto, E., Mahjour, B., Tsang, W., Xue, X., Islam, T., & Chen, W. (2020). OpenMP 4.5 compiler optimization for GPU offloading. IBM Journal of Research and Development, 64(3/4), 14:1-14:11. https://doi.org/10.1147/jrd.2019.2962428 DOI
  7. Venkataramanaiah, S. K., Ma, Y., Yin, S., Nurvithadhi, E., Dasu, A., Cao, Y., & Seo, J. S. (2019). Automatic Compiler Based FPGA Accelerator for CNN Training. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1908.06724 DOI
  8. Willsey, M., Lee, V. T., Cheung, A., Bodik, R., & Ceze, L. (2018). Iterative Search for Reconfigurable Accelerator Blocks With a Compiler in the Loop. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 38(3), 407-418. https://doi.org/10.1109/tcad.2018.2878194 DOI
  9. Xing, Y., Liang, S., Sui, L., Jia, X., Qiu, J., Liu, X., Wang, Y., Shan, Y., & Wang, Y. (2019). DNNVM: End-to-End Compiler Leveraging Heterogeneous Optimizations on FPGA-Based CNN Accelerators. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 39(10), 2668-2681. https://doi.org/10.1109/tcad.2019.2930577 DOI
  10. Şuşu, A. E. (2020). A Vector-Length Agnostic Compiler for the Connex-S Accelerator with Scratchpad Memory. ACM Transactions on Embedded Computing Systems, 19(6), 1-30. https://doi.org/10.1145/3406536 DOI
Publication details
Journal
Computing Architectures and Data Systems
Volume
1 (2026)
Article number
cads20260002
License
CC BY 4.0