跳到主要导航 跳到搜索 跳到主要内容

Efficient policy evaluation by matrix sketching

  • Cheng Chen
  • , Weinan Zhang*
  • , Yong Yu
  • *此作品的通讯作者
  • Shanghai Jiao Tong University

科研成果: 期刊稿件文章同行评审

摘要

In the reinforcement learning, policy evaluation aims to predict long-term values of a state under a certain policy. Since high-dimensional representations become more and more common in the reinforcement learning, how to reduce the computational cost becomes a significant problem to the policy evaluation. Many recent works focus on adopting matrix sketching methods to accelerate least-square temporal difference (TD) algorithms and quasi-Newton temporal difference algorithms. Among these sketching methods, the truncated incremental SVD shows better performance because it is stable and efficient. However, the convergence properties of the incremental SVD is still open. In this paper, we first show that the conventional incremental SVD algorithms could have enormous approximation errors in the worst case. Then we propose a variant of incremental SVD with better theoretical guarantees by shrinking the singular values periodically. Moreover, we employ our improved incremental SVD to accelerate least-square TD and quasi-Newton TD algorithms. The experimental results verify the correctness and effectiveness of our methods.

源语言英语
文章编号165330
期刊Frontiers of Computer Science
16
5
DOI
出版状态已出版 - 10月 2022
已对外发布

指纹

探究 'Efficient policy evaluation by matrix sketching' 的科研主题。它们共同构成独一无二的指纹。

引用此