Pages that link to "Item:Q2810885"
From MaRDI portal
The following pages link to An emphatic approach to the problem of off-policy temporal-difference learning (Q2810885):
Displaying 9 items.
- Adaptive importance sampling for value function approximation in off-policy reinforcement learning (Q1784527) (← links)
- Off-policy temporal difference learning with distribution adaptation in fast mixing chains (Q1797759) (← links)
- Multi-agent reinforcement learning: a selective overview of theories and algorithms (Q2094040) (← links)
- On Generalized Bellman Equations and Temporal-Difference Learning (Q3305109) (← links)
- Statistical Inference for Online Decision Making via Stochastic Gradient Descent (Q4999148) (← links)
- Efficiently Breaking the Curse of Horizon in Off-Policy Evaluation with Double Reinforcement Learning (Q5060503) (← links)
- Gradient temporal-difference learning for off-policy evaluation using emphatic weightings (Q6146179) (← links)
- Distributed consensus-based multi-agent temporal-difference learning (Q6164031) (← links)
- Online Bootstrap Inference For Policy Evaluation In Reinforcement Learning (Q6185586) (← links)