This review examines resilient data pipelines for scientific computing. The organizing question is how pipelines should preserve provenance and scientific meaning through retries, partial writes, schema change, and infrastructure failure. Ten related scholarly sources are synthesized through a decision-centered framework spanning problem definition, mechanism, measurement, evaluation, implementation, and governance. The review does not invent experiments, pooled estimates, or unreported quantitative results. It instead evaluates the strength and transferability of the available evidence, with particular attention to declaring computational success while silent corruption invalidates downstream evidence. The resulting framework links technical or empirical performance to explicit use conditions and identifies tests that should precede wider adoption in long-running and distributed scientific workflows.
- Ahmad, Z., Nazir, B., & Umer, A. (2020). A fault‐tolerant workflow management system with Quality‐of‐Service‐aware scheduling for scientific workflows in cloud computing. International Journal of Communication Systems, 34(1). https://doi.org/10.1002/dac.4649 DOI
- Alaei, M., Khorsand, R., & Ramezanpour, M. (2020). An adaptive fault detector strategy for scientific workflow scheduling based on improved differential evolution algorithm in cloud. Applied Soft Computing, 99, 106895. https://doi.org/10.1016/j.asoc.2020.106895 DOI
- Alaie, Y. A., Shirvani, M. H., & Rahmani, A. M. (2022). A hybrid bi-objective scheduling algorithm for execution of scientific workflows on cloud platforms with execution time and reliability approach. The Journal of Supercomputing, 79(2), 1451-1503. https://doi.org/10.1007/s11227-022-04703-0 DOI
- Bala, A., & Chana, I. (2015). Autonomic fault tolerant scheduling approach for scientific workflows in Cloud computing. Concurrent Engineering, 23(1), 27-39. https://doi.org/10.1177/1063293x14567783 DOI
- Chen, W., Silva, R. F. D., Deelman, E., & Fahringer, T. (2015). Dynamic and Fault-Tolerant Clustering for Scientific Workflows. IEEE Transactions on Cloud Computing, 4(1), 49-62. https://doi.org/10.1109/tcc.2015.2427200 DOI
- Khaldi, M., Rebbah, M., Meftah, B., & Smail, O. (2019). Fault tolerance for a scientific workflow system in a Cloud computing environment. International Journal of Computers and Applications, 42(7), 705-714. https://doi.org/10.1080/1206212x.2019.1647651 DOI
- Li, C., Liu, J., Wang, M., & Luo, Y. (2022). Fault-tolerant scheduling and data placement for scientific workflow processing in geo-distributed clouds. Journal of Systems and Software, 187, 111227. https://doi.org/10.1016/j.jss.2022.111227 DOI
- Li, Z., Chang, V., Hu, H., Hu, H., Li, C., & Ge, J. (2021). Real-time and dynamic fault-tolerant scheduling for scientific workflows in clouds. Information Sciences, 568, 13-39. https://doi.org/10.1016/j.ins.2021.03.003 DOI
- Tang, X. (2021). Reliability-Aware Cost-Efficient Scientific Workflows Scheduling Strategy on Multi-Cloud Systems. IEEE Transactions on Cloud Computing, 10(4), 2909-2919. https://doi.org/10.1109/tcc.2021.3057422 DOI
- Zhu, X., Wang, J., Guo, H., Zhu, D., Yang, L. T., & Liu, L. (2016). Fault-Tolerant Scheduling for Real-Time Scientific Workflows with Elastic Resource Provisioning in Virtualized Clouds. IEEE Transactions on Parallel and Distributed Systems, 27(12), 3501-3517. https://doi.org/10.1109/tpds.2016.2543731 DOI
- Journal
- Computing Architectures and Data Systems
- Volume
- 1 (2026)
- Article number
- cads20260005
- License
- CC BY 4.0