Offline Reinforcement Learning With Behavior Value Regularization.

Longyang Huang, Botao Dong, Wei Xie, Weidong Zhang

IEEE Transactions on Cybernetics 2024 April 27

Offline reinforcement learning (offline RL) aims to find task-solving policies from prerecorded datasets without online environment interaction. It is unfortunate that extrapolation errors can cause over-optimistic Q -value estimates when learning with a fixed dataset, limiting the performance of the learned policy. To tackle this issue, this article proposes an offline actor-critic with behavior value regularization (OAC-BVR) method. In the policy evaluation stage, the difference between the Q -function and the value of the behavior policy is considered as the regularization term, driving the learned value function to approach the value of the behavior policy. The convergence of the proposed policy evaluation with behavior value regularization (PE-BVR) and the value function difference are analyzed, respectively. Compared with existing offline actor-critic methods, the proposed OAC-BVR method integrates the value of the behavior policy, thereby simultaneously alleviating over-optimistic Q -value estimates and reducing Q -function bias. Experimental results on the D4RL MuJoCo and Maze2d datasets demonstrate the validity of the proposed PE-BVR and the performance advantage of OAC-BVR over the state-of-the-art offline RL algorithms. The code of OAC-BVR is available at https://github.com/LongyangHuang/OAC-BVR.

Full text links

We have located links that may give you full text access.

Show additional links to paperHide additional links to paper

PubMed

Add to Saved Papers

Get 1-tap access

Related Resources

Lung ultrasound for diagnosis and management of ARDS.Marry R Smit, Paul H Mayo, Silvia MongodiIntensive Care Medicine 2024 April 25

Executive Summary: State-of-the-Art Review: Unintended Consequences: Risk of Opportunistic Infections Associated with Long-term Glucocorticoid Therapies in Adults.Daniel B Chastain et al.Clinical Infectious Diseases 2024 April 11

Autoimmune Hemolytic Anemias: Classifications, Pathophysiology, Diagnoses and Management.Melika Loriamini et al.International Journal of Molecular Sciences 2024 April 13

Clinical practice guidelines on the management of status epilepticus in adults: A systematic review.Luca Vignatelli et al.Epilepsia 2024 April 13

Should renin-angiotensin system inhibitors be held prior to major surgery?Matthieu LegrandBritish Journal of Anaesthesia 2024 May

Contrast-induced acute kidney injury: a review of definition, pathogenesis, risk factors, prevention and treatment.Yanyan Li, Junda WangBMC Nephrology 2024 April 23

For the best experience, use the Read mobile app

Get seemless 1-tap access through your institution/university

For the best experience, use the Read mobile app

All material on this website is protected by copyright, Copyright © 1994-2024 by WebMD LLC.
This website also contains material copyrighted by 3rd parties.

By using this service, you agree to our terms of use and privacy policy.

Your Privacy Choices

You can now claim free CME credits for this literature searchClaim now

Get seemless 1-tap access through your institution/university

For the best experience, use the Read mobile app

Offline Reinforcement Learning With Behavior Value Regularization.

Full text links

Related Resources

Trending Papers

For the best experience, use the Read mobile app