End-to-end autonomous underwater vehicle path following control method based on improved soft actor–critic for deep space exploration

IF 10.4 1区计算机科学 Q1 COMPUTER SCIENCE, INTERDISCIPLINARY APPLICATIONS

Journal of Industrial Information Integration Pub Date : 2025-02-15 DOI:10.1016/j.jii.2025.100792

Na Dong , Shoufu Liu , Andrew W.H. Ip , Kai Leung Yung , Zhongke Gao , Rongshun Juan , Yanhui Wang

{"title":"End-to-end autonomous underwater vehicle path following control method based on improved soft actor–critic for deep space exploration","authors":"Na Dong , Shoufu Liu , Andrew W.H. Ip , Kai Leung Yung , Zhongke Gao , Rongshun Juan , Yanhui Wang","doi":"10.1016/j.jii.2025.100792","DOIUrl":null,"url":null,"abstract":"<div><div>The vast extraterrestrial ocean is becoming a hotspot for deep space exploration of life in the future. Considering autonomous underwater vehicle (AUV) has a larger range of activities and greater flexibility, it plays an important role in extraterrestrial ocean research. To solve the problems in path following tasks of AUV, such as high training cost and poor exploration ability, an end-to-end AUV path following control method based on an improved soft actor–critic (SAC) algorithm is designed in this paper, leveraging the advancements in deep reinforcement learning (DRL) to enhance performance and efficiency. It uses sensor information to understand the environment and its state to output the policy to complete the adaptive action. Policies that consider long-term effects can be learned through continuous interaction with the environment, which is helpful in improving adaptability and enhancing the robustness of AUV control. A non-policy sampling method is designed to improve the utilization efficiency of experience transitions in the replay buffer, accelerate convergence, and enhance its stability. A reward function on the current position and heading angle of AUV is designed to avoid the situation of sparse reward leading to slow learning or ineffective learning of agents. In the meantime, we use the continuous action space instead of the discrete action space to make the real-time control of the AUV more accurate. Finally, it is tested on the gazebo simulation platform, and the results confirm that reinforcement learning is effective in AUV control, and the method proposed in this paper has faster and better following performance than traditional reinforcement learning methods.</div></div>","PeriodicalId":55975,"journal":{"name":"Journal of Industrial Information Integration","volume":"45 ","pages":"Article 100792"},"PeriodicalIF":10.4000,"publicationDate":"2025-02-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Journal of Industrial Information Integration","FirstCategoryId":"94","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S2452414X25000160","RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"COMPUTER SCIENCE, INTERDISCIPLINARY APPLICATIONS","Score":null,"Total":0}

引用次数: 0

Abstract

The vast extraterrestrial ocean is becoming a hotspot for deep space exploration of life in the future. Considering autonomous underwater vehicle (AUV) has a larger range of activities and greater flexibility, it plays an important role in extraterrestrial ocean research. To solve the problems in path following tasks of AUV, such as high training cost and poor exploration ability, an end-to-end AUV path following control method based on an improved soft actor–critic (SAC) algorithm is designed in this paper, leveraging the advancements in deep reinforcement learning (DRL) to enhance performance and efficiency. It uses sensor information to understand the environment and its state to output the policy to complete the adaptive action. Policies that consider long-term effects can be learned through continuous interaction with the environment, which is helpful in improving adaptability and enhancing the robustness of AUV control. A non-policy sampling method is designed to improve the utilization efficiency of experience transitions in the replay buffer, accelerate convergence, and enhance its stability. A reward function on the current position and heading angle of AUV is designed to avoid the situation of sparse reward leading to slow learning or ineffective learning of agents. In the meantime, we use the continuous action space instead of the discrete action space to make the real-time control of the AUV more accurate. Finally, it is tested on the gazebo simulation platform, and the results confirm that reinforcement learning is effective in AUV control, and the method proposed in this paper has faster and better following performance than traditional reinforcement learning methods.

查看原文本刊更多论文

基于改进软评价的深空探测端到端自主潜航器路径跟踪控制方法

广阔的地外海洋正在成为未来深空生命探索的热点。由于自主水下航行器（AUV）具有更大的活动范围和更大的灵活性，在地外海洋研究中发挥着重要作用。针对AUV路径跟踪任务中训练成本高、探索能力差等问题，本文利用深度强化学习（DRL）技术的进步，设计了一种基于改进软行为者评价（SAC）算法的端到端AUV路径跟踪控制方法，以提高其性能和效率。利用传感器信息了解环境及其状态，输出策略完成自适应动作。考虑长期影响的策略可以通过与环境的持续交互来学习，这有助于提高AUV控制的适应性和鲁棒性。为了提高重放缓冲区中经验转换的利用效率，加快收敛速度，增强其稳定性，设计了一种非策略采样方法。设计了AUV当前位置和航向角的奖励函数，避免了奖励稀疏导致智能体学习缓慢或学习无效的情况。同时，我们用连续动作空间代替离散动作空间，使水下机器人的实时控制更加精确。最后，在gazebo仿真平台上进行了测试，结果证实了强化学习在AUV控制中的有效性，与传统的强化学习方法相比，本文提出的方法具有更快、更好的跟随性能。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Journal of Industrial Information Integration Decision Sciences-Information Systems and Management

CiteScore

22.30

自引率

13.40%

发文量

100

期刊介绍： The Journal of Industrial Information Integration focuses on the industry's transition towards industrial integration and informatization, covering not only hardware and software but also information integration. It serves as a platform for promoting advances in industrial information integration, addressing challenges, issues, and solutions in an interdisciplinary forum for researchers, practitioners, and policy makers. The Journal of Industrial Information Integration welcomes papers on foundational, technical, and practical aspects of industrial information integration, emphasizing the complex and cross-disciplinary topics that arise in industrial integration. Techniques from mathematical science, computer science, computer engineering, electrical and electronic engineering, manufacturing engineering, and engineering management are crucial in this context.