[1]ZHAO Ruohan,WEI Wei,WANG Da,et al.Dynamic distance-constrained offline-to-online reinforcement learning[J].CAAI Transactions on Intelligent Systems,2026,21(4):1055-1065.[doi:10.11992/tis.202510017]
Copy

Dynamic distance-constrained offline-to-online reinforcement learning

References:
[1] 刘全, 翟建伟, 章宗长, 等. 深度强化学习综述[J]. 计算机学报, 2018, 41(1): 1-27 LIU Quan, ZHAI Jianwei, ZHANG Zongchang, et al. A survey on deep reinforcement learning[J]. Chinese journal of computers, 2018, 41(1): 1-27
[2] 王硕汝, 牛温佳, 童恩栋, 等. 强化学习离线策略评估研究综述[J]. 计算机学报, 2022, 45(9): 1926-1945 WANG Shuoru, NIU Wenjia, TONG Endong, et al. Research on off-policy evaluation in reinforcement learning: a survey[J]. Chinese journal of computers, 2022, 45(9): 1926-1945
[3] TANG Shengpu, WIENS J. Model selection for offline reinforcement learning: practical considerations for healthcare settings[EB/OL]. (2021-07-23)[2025-10-16]. https://arxiv.org/abs/2107.11003.
[4] 刘健, 顾扬, 程玉虎, 等. 基于多智能体强化学习的乳腺癌致病基因预测[J]. 自动化学报, 2022, 48(5): 1246-1258 LIU Jian, GU Yang, CHENG Yuhu, et al. Prediction of breast cancer pathogenic genes based on multi-agent reinforcement learning[J]. Acta automatica sinica, 2022, 48(5): 1246-1258
[5] 吴晓光, 刘绍维, 杨磊, 等. 基于深度强化学习的双足机器人斜坡步态控制方法[J]. 自动化学报, 2021, 47(8): 1976-1987 WU Xiaoguang, LIU Shaowei, YANG Lei, et al. A gait control method for biped robot on slope based on deep reinforcement learning[J]. Acta automatica sinica, 2021, 47(8): 1976-1987
[6] XIAO Teng, WANG Donglin. A general offline reinforcement learning framework for interactive recommendation[J]. Proceedings of the AAAI conference on artificial intelligence, 2021, 35(5): 4512-4520
[7] GAO Chongming, WANG Shiqi, LI Shijun, et al. CIRS: bursting filter bubbles by counterfactual interactive recommender system[J]. ACM transactions on information systems, 2024, 42(1): 1-27
[8] NAIR A, GUPTA A, DALAL M, et al. AWAC: accelerating online reinforcement learning with offline datasets[EB/OL]. (2020-06-17)[2025-09-11]. https://arxiv.org/abs/2006.09359.
[9] MAO Yihuan, WANG Chao, WANG Bin, et al. MOORe: model-based offline-to-online reinforcement learning[EB/OL]. (2022-01-23)[2025-09-11]. https://arxiv.org/abs/2201.10070.
[10] BEESON A, MONTANA G. Improving TD3-BC: relaxed policy constraint for offline learning and stable online fine-tuning[EB/OL]. (2022-11-21)[2025-09-11]. https://arxiv.org/abs/2211.11802.
[11] GUO Siyuan, SUN Yanchao, HU Jifeng, et al. A simple unified uncertainty-guided framework for offline-to-online reinforcement learning[EB/OL]. (2023-06-13)[2025-09-11]. https://arxiv.org/abs/2306.07541.
[12] ZHENG Han, LUO Xufang, WEI Pengfei, et al. Adaptive policy learning for offline-to-online reinforcement learning[J]. Proceedings of the AAAI conference on artificial intelligence, 2023, 37(9): 11372-11380
[13] LUO Yicheng, KAY J, GREFENSTETTE E, et al. Finetuning from offline reinforcement learning: challenges, trade-offs and practical solutions[EB/OL]. (2023-03-30)[2025-09-11]. https://arxiv.org/abs/2303.17396.
[14] LI Jianxiong, HU Xiao, XU Haoran, et al. PROTO: iterative policy regularized offline-to-online reinforcement learning[EB/OL]. (2023-05-25)[2025-09-11]. https://arxiv.org/abs/2305.15669.
[15] HU Hao, YANG Yiqin, YE Jianing, et al. Bayesian design principles for offline-to-online reinforcement learning[EB/OL]. (2024-05-31)[2025-09-11]. https://arxiv.org/abs/2405.20984.
[16] KONG Rui, WU Chenyang, GAO Chenxiao, et al. Efficient and stable offline-to-online reinforcement learning via continual policy revitalization[C]//International Joint Conference on Artificial Intelligence. Jeju: IJCAI, 2024.
[17] SUTTON R S, BARTO A G. Reinforcement Learning[M]. Cambridge: MIT Press, 1998: 9-11.
[18] KUMAR A, ZHOU A, TUCKER G, et al. Conservative Q-learning for offline reinforcement learning[C]//Advances in Neural Information Processing Systems. New York: Curran Associates Inc. , 2020: 1179-1191.
[19] FUJIMOTO S, MEGER D, PRECUP D. Off-policy deep reinforcement learning without exploration[C]//Proceedings of the 36th International Conference on Machine Learning. Long Beach: PMLR, 2019: 2052-2062.
[20] FIGUEIREDO PRUDENCIO R, MAXIMO M R O A, COLOMBINI E L. A survey on offline reinforcement learning: taxonomy, review, and open problems[J]. IEEE transactions on neural networks and learning systems, 2024, 35(8): 10237-10257
[21] LI Jianxiong, ZHAN Xianyuan, XU Haoran, et al. When data geometry meets deep function: generalizing offline reinforcement learning[EB/OL]. (2022-05-23)[2025-09-16]. https://arxiv.org/abs/2205.11027.
[22] FUJIMOTO S, VAN HOOF H, MEGER D. Addressing function approximation error in actor-critic methods[C]//International Conference on Machine Learning. Stockholm: PMLR, 2018: 1587-1596.
[23] ZHENG Qinqing, ZHANG A, GROVER A. Online decision Transformer[C]//Proceedings of the 39th International Conference on Machine Learning. Baltimore: PMLR, 2022: 27042-27059.
[24] KOSTRIKOV I, NAIR A, LEVINE S. Offline reinforcement learning with implicit Q-learning[EB/OL]. (2021-10-12)[2025-09-16]. https://arxiv.org/abs/2110.06169.
[25] ZHANG Haichao, XU W, YU Haonan. Policy expansion for bridging offline-to-online reinforcement learning[EB/OL]. (2023-02-02)[2025-09-16]. https://arxiv.org/abs/2302.00935.
[26] FINN C, KUMAR A, LEVINE S, et al. Cal-QL: calibrated offline RL pre-training for efficient online fine-tuning[C]//Advances in Neural Information Processing Systems 36. New Orleans: Neural Information Processing Systems Foundation, Inc, 2023: 62244-62269.
[27] GUO Siyuan, ZOU Lixin, CHEN Hechang, et al. Sample efficient offline-to-online reinforcement learning[J]. IEEE transactions on knowledge and data engineering, 2024, 36(3): 1299-1310
[28] FU J, KUMAR A, NACHUM O, et al. D4RL: datasets for deep data-driven reinforcement learning[EB/OL]. (2020-04-15)[2025-09-16]. https://arxiv.org/abs/2004.07219.
[29] FUJIMOTO S, GU S S. A minimalist approach to offlinereinforcement learning[C]//Advances in Neural Infor-mation Processing Systems. Virtual Conference: CurranAssociates Inc. , 2021: 20132-20145.
[30] LEE S, SEO Y, LEE K, et al. Offline-to-online reinforcement learning via balanced replay and pessimistic Q-ensemble[EB/OL]. (2021-07-15)[2025-09-16]. https://arxiv.org/abs/2107.00591.
[31] CHEN Hao, GAO Jiawei, HUANG Gao, et al. Train once, get a family: state-adaptive balances for offline-to-online reinforcement learning[C]//Advances in Neural Information Processing Systems 36. New Orleans: Neural Information Processing Systems Foundation, Inc. , 2023: 47081-47104.
[32] LIU Xuhui, LIU Tianshuo, JIANG Shengyi, et al. Energy-guided diffusion sampling for offline-to-online reinforcement learning[EB/OL]. (2024-07-17)[2025-09-16]. https://arxiv.org/abs/2407.12448.
[33] HE Longxiang, YE Deheng, TAN Junbo, et al. Robust policy expansion for offline-to-online RL under diverse data corruption[EB/OL]. (2025-09-29)[2025-10-03]. https://arxiv.org/abs/2509.24748.
[34] HUANG Xiao, LIU Xu, ZHANG Enze, et al. Offline-to-online reinforcement learning with classifier-free diffusion generation[EB/OL]. (2025-08-09)[2025-09-16]. https://arxiv.org/abs/2508.06806.
[35] HUANG Shengjun, LUO Qinwen, WANG Yewen, et al. Optimistic critic reconstruction and constrained fine-tuning for general offline-to-online RL[C]//Advances in Neural Information Processing Systems 37. Vancouver: Neural Information Processing Systems Foundation, Inc. , 2024: 108167-108207.
[36] CHEN Siqi, ZHAO Jianing, ZHAO Kai, et al. ANOTO: improving automated negotiation via offline-to-online reinforcement learning[C]//Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Systems. London: ACM, 2024: 2195-2197.
Similar References:

Memo

-

Last Update: 1900-01-01

Copyright © CAAI Transactions on Intelligent Systems