[1]ZHAO Ruohan,WEI Wei,WANG Da,et al.Dynamic distance-constrained offline-to-online reinforcement learning[J].CAAI Transactions on Intelligent Systems,2026,21(4):1055-1065.[doi:10.11992/tis.202510017]
Copy
CAAI Transactions on Intelligent Systems[ISSN 1673-4785/CN 23-1538/TP] Volume:
21
Number of periods:
2026 4
Page number:
1055-1065
Column:
人工智能院长论坛
Public date:
2026-07-05
- Title:
-
Dynamic distance-constrained offline-to-online reinforcement learning
- Author(s):
-
ZHAO Ruohan; WEI Wei; WANG Da; HUANG Li
-
School of Computer Science and Information Technology, Shanxi University, Taiyuan 030006, China
-
- Keywords:
-
machine learning; reinforcement learning; offline-to-online reinforcement learning; Markov processes; policy optimization; distance constraint; policy transfer; constrained optimization
- CLC:
-
TP181
- DOI:
-
10.11992/tis.202510017
- Abstract:
-
Offline-to-online reinforcement learning enables rapid policy improvement with limited online interactions. However, in the early stage of fine-tuning, the policy distribution often deviates from the support region of offline data, leading to unstable value estimation and restricting performance gains. Existing methods typically rely on uniform conservative constraints, treating offline data as boundary conditions, but they overlook sample quality and distributional differences. As a result, policy updates are simultaneously over-constrained and prone to unstable estimation outside the offline distribution. To address this issue, we propose a dynamic distance-constrained algorithm for offline-to-online reinforcement learning. Our method employs a state-conditioned distance function to constrain the policy within the support region of offline data and gradually relaxes the constraint toward high-quality data regions as online interactions progress, thereby achieving a smooth transition from offline to online learning. Experimental results demonstrate that, compared with mainstream baselines, the proposed method achieves stable and effective performance improvements across multiple simulated tasks.