Learning in Observable POMDPs, without Computationally Intractable Oracles

Golowich, Noah
Moitra, Ankur
Rohatgi, Dhruv

Publication date

June 2022

Abstract

Much of reinforcement learning theory is built on top of oracles that are computationally hard to implement. Specifically for learning near-optimal policies in Partially Observable Markov Decision Processes (POMDPs), existing algorithms either need to make strong assumptions about the model dynamics (e.g. deterministic transitions) or assume access to an oracle for solving a hard optimistic planning or estimation problem as a subroutine. In this work we develop the first oracle-free learning algorithm for POMDPs under reasonable assumptions. Specifically, we give a quasipolynomial-time end-to-end algorithm for learning in "observable" POMDPs, where observability is the assumption that well-separated distributions over states induce well-sep...

Extracted data

We use cookies to provide a better user experience.

Data Protection

Learning in Observable POMDPs, without Computationally Intractable Oracles

Abstract

Extracted data

Learning in Observable POMDPs, without Computationally Intractable Oracles

Abstract

Extracted data

Related items

Related items