[1]赵若涵,魏巍,王达,等.动态距离约束的离线到在线强化学习[J].智能系统学报,2026,21(4):1055-1065.[doi:10.11992/tis.202510017]
 ZHAO Ruohan,WEI Wei,WANG Da,et al.Dynamic distance-constrained offline-to-online reinforcement learning[J].CAAI Transactions on Intelligent Systems,2026,21(4):1055-1065.[doi:10.11992/tis.202510017]
点击复制

动态距离约束的离线到在线强化学习

参考文献/References:
[1] 刘全, 翟建伟, 章宗长, 等. 深度强化学习综述[J]. 计算机学报, 2018, 41(1): 1-27 LIU Quan, ZHAI Jianwei, ZHANG Zongchang, et al. A survey on deep reinforcement learning[J]. Chinese journal of computers, 2018, 41(1): 1-27
[2] 王硕汝, 牛温佳, 童恩栋, 等. 强化学习离线策略评估研究综述[J]. 计算机学报, 2022, 45(9): 1926-1945 WANG Shuoru, NIU Wenjia, TONG Endong, et al. Research on off-policy evaluation in reinforcement learning: a survey[J]. Chinese journal of computers, 2022, 45(9): 1926-1945
[3] TANG Shengpu, WIENS J. Model selection for offline reinforcement learning: practical considerations for healthcare settings[EB/OL]. (2021-07-23)[2025-10-16]. https://arxiv.org/abs/2107.11003.
[4] 刘健, 顾扬, 程玉虎, 等. 基于多智能体强化学习的乳腺癌致病基因预测[J]. 自动化学报, 2022, 48(5): 1246-1258 LIU Jian, GU Yang, CHENG Yuhu, et al. Prediction of breast cancer pathogenic genes based on multi-agent reinforcement learning[J]. Acta automatica sinica, 2022, 48(5): 1246-1258
[5] 吴晓光, 刘绍维, 杨磊, 等. 基于深度强化学习的双足机器人斜坡步态控制方法[J]. 自动化学报, 2021, 47(8): 1976-1987 WU Xiaoguang, LIU Shaowei, YANG Lei, et al. A gait control method for biped robot on slope based on deep reinforcement learning[J]. Acta automatica sinica, 2021, 47(8): 1976-1987
[6] XIAO Teng, WANG Donglin. A general offline reinforcement learning framework for interactive recommendation[J]. Proceedings of the AAAI conference on artificial intelligence, 2021, 35(5): 4512-4520
[7] GAO Chongming, WANG Shiqi, LI Shijun, et al. CIRS: bursting filter bubbles by counterfactual interactive recommender system[J]. ACM transactions on information systems, 2024, 42(1): 1-27
[8] NAIR A, GUPTA A, DALAL M, et al. AWAC: accelerating online reinforcement learning with offline datasets[EB/OL]. (2020-06-17)[2025-09-11]. https://arxiv.org/abs/2006.09359.
[9] MAO Yihuan, WANG Chao, WANG Bin, et al. MOORe: model-based offline-to-online reinforcement learning[EB/OL]. (2022-01-23)[2025-09-11]. https://arxiv.org/abs/2201.10070.
[10] BEESON A, MONTANA G. Improving TD3-BC: relaxed policy constraint for offline learning and stable online fine-tuning[EB/OL]. (2022-11-21)[2025-09-11]. https://arxiv.org/abs/2211.11802.
[11] GUO Siyuan, SUN Yanchao, HU Jifeng, et al. A simple unified uncertainty-guided framework for offline-to-online reinforcement learning[EB/OL]. (2023-06-13)[2025-09-11]. https://arxiv.org/abs/2306.07541.
[12] ZHENG Han, LUO Xufang, WEI Pengfei, et al. Adaptive policy learning for offline-to-online reinforcement learning[J]. Proceedings of the AAAI conference on artificial intelligence, 2023, 37(9): 11372-11380
[13] LUO Yicheng, KAY J, GREFENSTETTE E, et al. Finetuning from offline reinforcement learning: challenges, trade-offs and practical solutions[EB/OL]. (2023-03-30)[2025-09-11]. https://arxiv.org/abs/2303.17396.
[14] LI Jianxiong, HU Xiao, XU Haoran, et al. PROTO: iterative policy regularized offline-to-online reinforcement learning[EB/OL]. (2023-05-25)[2025-09-11]. https://arxiv.org/abs/2305.15669.
[15] HU Hao, YANG Yiqin, YE Jianing, et al. Bayesian design principles for offline-to-online reinforcement learning[EB/OL]. (2024-05-31)[2025-09-11]. https://arxiv.org/abs/2405.20984.
[16] KONG Rui, WU Chenyang, GAO Chenxiao, et al. Efficient and stable offline-to-online reinforcement learning via continual policy revitalization[C]//International Joint Conference on Artificial Intelligence. Jeju: IJCAI, 2024.
[17] SUTTON R S, BARTO A G. Reinforcement Learning[M]. Cambridge: MIT Press, 1998: 9-11.
[18] KUMAR A, ZHOU A, TUCKER G, et al. Conservative Q-learning for offline reinforcement learning[C]//Advances in Neural Information Processing Systems. New York: Curran Associates Inc. , 2020: 1179-1191.
[19] FUJIMOTO S, MEGER D, PRECUP D. Off-policy deep reinforcement learning without exploration[C]//Proceedings of the 36th International Conference on Machine Learning. Long Beach: PMLR, 2019: 2052-2062.
[20] FIGUEIREDO PRUDENCIO R, MAXIMO M R O A, COLOMBINI E L. A survey on offline reinforcement learning: taxonomy, review, and open problems[J]. IEEE transactions on neural networks and learning systems, 2024, 35(8): 10237-10257
[21] LI Jianxiong, ZHAN Xianyuan, XU Haoran, et al. When data geometry meets deep function: generalizing offline reinforcement learning[EB/OL]. (2022-05-23)[2025-09-16]. https://arxiv.org/abs/2205.11027.
[22] FUJIMOTO S, VAN HOOF H, MEGER D. Addressing function approximation error in actor-critic methods[C]//International Conference on Machine Learning. Stockholm: PMLR, 2018: 1587-1596.
[23] ZHENG Qinqing, ZHANG A, GROVER A. Online decision Transformer[C]//Proceedings of the 39th International Conference on Machine Learning. Baltimore: PMLR, 2022: 27042-27059.
[24] KOSTRIKOV I, NAIR A, LEVINE S. Offline reinforcement learning with implicit Q-learning[EB/OL]. (2021-10-12)[2025-09-16]. https://arxiv.org/abs/2110.06169.
[25] ZHANG Haichao, XU W, YU Haonan. Policy expansion for bridging offline-to-online reinforcement learning[EB/OL]. (2023-02-02)[2025-09-16]. https://arxiv.org/abs/2302.00935.
[26] FINN C, KUMAR A, LEVINE S, et al. Cal-QL: calibrated offline RL pre-training for efficient online fine-tuning[C]//Advances in Neural Information Processing Systems 36. New Orleans: Neural Information Processing Systems Foundation, Inc, 2023: 62244-62269.
[27] GUO Siyuan, ZOU Lixin, CHEN Hechang, et al. Sample efficient offline-to-online reinforcement learning[J]. IEEE transactions on knowledge and data engineering, 2024, 36(3): 1299-1310
[28] FU J, KUMAR A, NACHUM O, et al. D4RL: datasets for deep data-driven reinforcement learning[EB/OL]. (2020-04-15)[2025-09-16]. https://arxiv.org/abs/2004.07219.
[29] FUJIMOTO S, GU S S. A minimalist approach to offlinereinforcement learning[C]//Advances in Neural Infor-mation Processing Systems. Virtual Conference: CurranAssociates Inc. , 2021: 20132-20145.
[30] LEE S, SEO Y, LEE K, et al. Offline-to-online reinforcement learning via balanced replay and pessimistic Q-ensemble[EB/OL]. (2021-07-15)[2025-09-16]. https://arxiv.org/abs/2107.00591.
[31] CHEN Hao, GAO Jiawei, HUANG Gao, et al. Train once, get a family: state-adaptive balances for offline-to-online reinforcement learning[C]//Advances in Neural Information Processing Systems 36. New Orleans: Neural Information Processing Systems Foundation, Inc. , 2023: 47081-47104.
[32] LIU Xuhui, LIU Tianshuo, JIANG Shengyi, et al. Energy-guided diffusion sampling for offline-to-online reinforcement learning[EB/OL]. (2024-07-17)[2025-09-16]. https://arxiv.org/abs/2407.12448.
[33] HE Longxiang, YE Deheng, TAN Junbo, et al. Robust policy expansion for offline-to-online RL under diverse data corruption[EB/OL]. (2025-09-29)[2025-10-03]. https://arxiv.org/abs/2509.24748.
[34] HUANG Xiao, LIU Xu, ZHANG Enze, et al. Offline-to-online reinforcement learning with classifier-free diffusion generation[EB/OL]. (2025-08-09)[2025-09-16]. https://arxiv.org/abs/2508.06806.
[35] HUANG Shengjun, LUO Qinwen, WANG Yewen, et al. Optimistic critic reconstruction and constrained fine-tuning for general offline-to-online RL[C]//Advances in Neural Information Processing Systems 37. Vancouver: Neural Information Processing Systems Foundation, Inc. , 2024: 108167-108207.
[36] CHEN Siqi, ZHAO Jianing, ZHAO Kai, et al. ANOTO: improving automated negotiation via offline-to-online reinforcement learning[C]//Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Systems. London: ACM, 2024: 2195-2197.
相似文献/References:
[1]叶志飞,文益民,吕宝粮.不平衡分类问题研究综述[J].智能系统学报,2009,4(2):148.
 YE Zhi-fei,WEN Yi-min,LU Bao-liang.A survey of imbalanced pattern classification problems[J].CAAI Transactions on Intelligent Systems,2009,4():148.
[2]刘奕群,张 敏,马少平.基于非内容信息的网络关键资源有效定位[J].智能系统学报,2007,2(1):45.
 LIU Yi-qun,ZHANG Min,MA Shao-ping.Web key resource page selection based on non-content inf o rmation[J].CAAI Transactions on Intelligent Systems,2007,2():45.
[3]马世龙,眭跃飞,许 可.优先归纳逻辑程序的极限行为[J].智能系统学报,2007,2(4):9.
 MA Shi-long,SUI Yue-fei,XU Ke.Limit behavior of prioritized inductive logic programs[J].CAAI Transactions on Intelligent Systems,2007,2():9.
[4]连传强,徐昕,吴军,等.面向资源分配问题的Q-CF多智能体强化学习[J].智能系统学报,2011,6(2):95.
 LIAN Chuanqiang,XU Xin,WU Jun,et al.Q-CF multiAgent reinforcement learningfor resource allocation problems[J].CAAI Transactions on Intelligent Systems,2011,6():95.
[5]姚伏天,钱沄涛.高斯过程及其在高光谱图像分类中的应用[J].智能系统学报,2011,6(5):396.
 YAO Futian,QIAN Yuntao.Gaussian process and its applications in hyperspectral image classification[J].CAAI Transactions on Intelligent Systems,2011,6():396.
[6]文益民,强保华,范志刚.概念漂移数据流分类研究综述[J].智能系统学报,2013,8(2):95.[doi:10.3969/j.issn.1673-4785.201208012]
 WEN Yimin,QIANG Baohua,FAN Zhigang.A survey of the classification of data streams with concept drift[J].CAAI Transactions on Intelligent Systems,2013,8():95.[doi:10.3969/j.issn.1673-4785.201208012]
[7]杨成东,邓廷权.综合属性选择和删除的属性约简方法[J].智能系统学报,2013,8(2):183.[doi:10.3969/j.issn.1673-4785.201209056]
 YANG Chengdong,DENG Tingquan.An approach to attribute reduction combining attribute selection and deletion[J].CAAI Transactions on Intelligent Systems,2013,8():183.[doi:10.3969/j.issn.1673-4785.201209056]
[8]胡小生,钟勇.基于加权聚类质心的SVM不平衡分类方法[J].智能系统学报,2013,8(3):261.
 HU Xiaosheng,ZHONG Yong.Support vector machine imbalanced data classification based on weighted clustering centroid[J].CAAI Transactions on Intelligent Systems,2013,8():261.
[9]丁科,谭营.GPU通用计算及其在计算智能领域的应用[J].智能系统学报,2015,10(1):1.[doi:10.3969/j.issn.1673-4785.201403072]
 DING Ke,TAN Ying.A review on general purpose computing on GPUs and its applications in computational intelligence[J].CAAI Transactions on Intelligent Systems,2015,10():1.[doi:10.3969/j.issn.1673-4785.201403072]
[10]孔庆超,毛文吉,张育浩.社交网站中用户评论行为预测[J].智能系统学报,2015,10(3):349.[doi:10.3969/j.issn.1673-4785.201403019]
 KONG Qingchao,MAO Wenji,ZHANG Yuhao.User comment behavior prediction in social networking sites[J].CAAI Transactions on Intelligent Systems,2015,10():349.[doi:10.3969/j.issn.1673-4785.201403019]
[11]周文吉,俞扬.分层强化学习综述[J].智能系统学报,2017,12(5):590.[doi:10.11992/tis.201706031]
 ZHOU Wenji,YU Yang.Summarize of hierarchical reinforcement learning[J].CAAI Transactions on Intelligent Systems,2017,12():590.[doi:10.11992/tis.201706031]
[12]殷昌盛,杨若鹏,朱巍,等.多智能体分层强化学习综述[J].智能系统学报,2020,15(4):646.[doi:10.11992/tis.201909027]
 YIN Changsheng,YANG Ruopeng,ZHU Wei,et al.A survey on multi-agent hierarchical reinforcement learning[J].CAAI Transactions on Intelligent Systems,2020,15():646.[doi:10.11992/tis.201909027]
[13]杨瑞,严江鹏,李秀.强化学习稀疏奖励算法研究——理论与实验[J].智能系统学报,2020,15(5):888.[doi:10.11992/tis.202003031]
 YANG Rui,YAN Jiangpeng,LI Xiu.Survey of sparse reward algorithms in reinforcement learning — theory and experiment[J].CAAI Transactions on Intelligent Systems,2020,15():888.[doi:10.11992/tis.202003031]
[14]高春艳,刘琦,李满宏,等.面向复杂环境的机器人触觉感知算法研究综述[J].智能系统学报,2026,21(4):834.[doi:10.11992/tis.202510008]
 GAO Chunyan,LIU Qi,LI Manhong,et al.Review of robot tactile perception algorithms for complex environments[J].CAAI Transactions on Intelligent Systems,2026,21():834.[doi:10.11992/tis.202510008]

备注/Memo

收稿日期:2025-10-16。
基金项目:国家自然科学基金项目(62276160);山西省基础研究计划(202203021211294).
作者简介:赵若涵,硕士研究生,主要研究方向为强化学习。E-mail:zruohan2023@163.com。;魏巍,教授,博士生导师,山西大学计算机与信息技术学院(大数据学院)副院长,主要研究方向为数据挖掘、机器学习与具身智能。主持和参与国家重点研发计划项目、国家自然科学基金重点项目、国家自然科学基金面上项目、山西省自然科学基金项目20余项。发表学术论文40余篇。E-mail:weiwei@sxu.edu.cn。;王达,讲师,博士,主要研究方向为强化学习和具身智能,获国家发明专利授权2项,发表学术论文11篇。E-mail:wanda@sxu.edu.cn。
通讯作者:魏巍. E-mail:weiwei@sxu.edu.cn

更新日期/Last Update: 1900-01-01
Copyright © 《 智能系统学报》 编辑部
地址:(150001)黑龙江省哈尔滨市南岗区南通大街145-1号楼 电话:0451- 82534001、82518134 邮箱:tis@vip.sina.com