This review examines safe exploration in embodied reinforcement learning. The organizing question is how agents can learn informative behaviour while respecting physical, social, and operational constraints. Ten related scholarly sources are synthesized through a decision-centered framework spanning problem definition, mechanism, measurement, evaluation, implementation, and governance. The review does not invent experiments, pooled estimates, or unreported quantitative results. It instead evaluates the strength and transferability of the available evidence, with particular attention to treating simulated constraint satisfaction as proof of physical-world safety. The resulting framework links technical or empirical performance to explicit use conditions and identifies tests that should precede wider adoption in robotics and autonomous embodied systems.
- Alshiekh, M., Bloem, R., Ehlers, R., Könighofer, B., Niekum, S., & Topcu, U. (2018). Safe Reinforcement Learning via Shielding. Scholarly publication. https://doi.org/10.1609/aaai.v32i1.11797 DOI
- Basso, R., Kulcsár, B., Sanchez-Diaz, I., & Qu, X. (2021). Dynamic stochastic electric vehicle routing with safe reinforcement learning. Transportation Research Part E Logistics and Transportation Review, 157, 102496. https://doi.org/10.1016/j.tre.2021.102496 DOI
- Brunke, L., Greeff, M., Hall, A. W., Yuan, Z., Zhou, S., Panerati, J., & Schoellig, A. P. (2022). Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning. Annual Review of Control Robotics and Autonomous Systems, 5(1), 411-444. https://doi.org/10.1146/annurev-control-042920-020211 DOI
- GarcíaJavier, & FernándezFernando (2015). A comprehensive survey on safe reinforcement learning. Journal of Machine Learning Research. https://doi.org/10.5555/2789272.2886795 DOI
- Gu, S., Yang, L., Du, Y., Chen, G., Walter, F., Wang, J., & Knoll, A. (2024). A Review of Safe Reinforcement Learning: Methods, Theories, and Applications. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12), 11216-11235. https://doi.org/10.1109/tpami.2024.3457538 DOI
- Marvi, Z., & Kiumarsi, B. (2020). Safe reinforcement learning: A control barrier function optimization approach. International Journal of Robust and Nonlinear Control, 31(6), 1923-1940. https://doi.org/10.1002/rnc.5132 DOI
- Thananjeyan, B., Balakrishna, A., Nair, S., Luo, M., Srinivasan, K., Hwang, M., Gonzalez, J. E., Ibarz, J., Finn, C., & Goldberg, K. (2021). Recovery RL: Safe Reinforcement Learning With Learned Recovery Zones. IEEE Robotics and Automation Letters, 6(3), 4915-4922. https://doi.org/10.1109/lra.2021.3070252 DOI
- Yang, Y., Vamvoudakis, K. G., & Modares, H. (2020). Safe reinforcement learning for dynamical games. International Journal of Robust and Nonlinear Control, 30(9), 3706-3726. https://doi.org/10.1002/rnc.4962 DOI
- Zanon, M., & Gros, S. (2020). Safe Reinforcement Learning Using Robust MPC. IEEE Transactions on Automatic Control, 66(8), 3638-3652. https://doi.org/10.1109/tac.2020.3024161 DOI
- Zhang, L., Zhang, R., Wu, T., Weng, R., Han, M., & Zhao, Y. (2021). Safe Reinforcement Learning With Stability Guarantee for Motion Planning of Autonomous Vehicles. IEEE Transactions on Neural Networks and Learning Systems, 32(12), 5435-5444. https://doi.org/10.1109/tnnls.2021.3084685 DOI
- Journal
- Advances in Adaptive Intelligence
- Volume
- 1 (2026)
- Article number
- aai20260002
- License
- CC BY 4.0