[1]赵若涵,魏巍,王达,等.动态距离约束的离线到在线强化学习[J].智能系统学报,2026,21(4):1055-1065.[doi:10.11992/tis.202510017]
ZHAO Ruohan,WEI Wei,WANG Da,et al.Dynamic distance-constrained offline-to-online reinforcement learning[J].CAAI Transactions on Intelligent Systems,2026,21(4):1055-1065.[doi:10.11992/tis.202510017]
点击复制
《智能系统学报》[ISSN 1673-4785/CN 23-1538/TP] 卷:
21
期数:
2026年第4期
页码:
1055-1065
栏目:
人工智能院长论坛
出版日期:
2026-07-05
- Title:
-
Dynamic distance-constrained offline-to-online reinforcement learning
- 作者:
-
赵若涵, 魏巍, 王达, 黄利
-
山西大学 计算机与信息技术学院, 山西 太原 030006
- Author(s):
-
ZHAO Ruohan, WEI Wei, WANG Da, HUANG Li
-
School of Computer Science and Information Technology, Shanxi University, Taiyuan 030006, China
-
- 关键词:
-
机器学习; 强化学习; 离线到在线强化学习; 马尔可夫过程; 策略优化; 距离约束; 策略迁移; 约束优化
- Keywords:
-
machine learning; reinforcement learning; offline-to-online reinforcement learning; Markov processes; policy optimization; distance constraint; policy transfer; constrained optimization
- 分类号:
-
TP181
- DOI:
-
10.11992/tis.202510017
- 摘要:
-
离线到在线强化学习能够在有限的在线交互下快速提升策略性能,但在微调初期,策略分布往往偏离离线数据支持区域,导致估值不稳并限制性能提升。现有方法多依赖统一的保守约束,将离线数据视作边界条件,却忽视了样本质量与分布差异,从而在更新时既过度约束,又难以避免分布外的不稳定估值。为此,本文提出一种动态距离约束的离线到在线强化学习算法。该方法基于状态条件距离函数,将策略约束在离线数据支持区域内,并随着在线交互的进行,动态地将约束放宽至高质量数据区域,从而实现离线到在线的平滑迁移。实验结果表明,与现有主流方法相比,所提出的方法可以在多个模拟任务上实现稳定有效的性能改进。
- Abstract:
-
Offline-to-online reinforcement learning enables rapid policy improvement with limited online interactions. However, in the early stage of fine-tuning, the policy distribution often deviates from the support region of offline data, leading to unstable value estimation and restricting performance gains. Existing methods typically rely on uniform conservative constraints, treating offline data as boundary conditions, but they overlook sample quality and distributional differences. As a result, policy updates are simultaneously over-constrained and prone to unstable estimation outside the offline distribution. To address this issue, we propose a dynamic distance-constrained algorithm for offline-to-online reinforcement learning. Our method employs a state-conditioned distance function to constrain the policy within the support region of offline data and gradually relaxes the constraint toward high-quality data regions as online interactions progress, thereby achieving a smooth transition from offline to online learning. Experimental results demonstrate that, compared with mainstream baselines, the proposed method achieves stable and effective performance improvements across multiple simulated tasks.
备注/Memo
收稿日期:2025-10-16。
基金项目:国家自然科学基金项目(62276160);山西省基础研究计划(202203021211294).
作者简介:赵若涵,硕士研究生,主要研究方向为强化学习。E-mail:zruohan2023@163.com。;魏巍,教授,博士生导师,山西大学计算机与信息技术学院(大数据学院)副院长,主要研究方向为数据挖掘、机器学习与具身智能。主持和参与国家重点研发计划项目、国家自然科学基金重点项目、国家自然科学基金面上项目、山西省自然科学基金项目20余项。发表学术论文40余篇。E-mail:weiwei@sxu.edu.cn。;王达,讲师,博士,主要研究方向为强化学习和具身智能,获国家发明专利授权2项,发表学术论文11篇。E-mail:wanda@sxu.edu.cn。
通讯作者:魏巍. E-mail:weiwei@sxu.edu.cn
更新日期/Last Update:
1900-01-01