Pages that link to "Item:Q4974645"
From MaRDI portal
The following pages link to Convergence Results for Some Temporal Difference Methods Based on Least Squares (Q4974645):
Displaying 13 items.
- Temporal difference-based policy iteration for optimal control of stochastic systems (Q467477) (← links)
- Approximate dynamic programming via direct search in the space of value function approximations (Q713118) (← links)
- Proximal algorithms and temporal difference methods for solving fixed point problems (Q721950) (← links)
- Regularized feature selection in reinforcement learning (Q747290) (← links)
- Concentration bounds for temporal difference learning with linear function approximation: the case of batch data and uniform sampling (Q2051259) (← links)
- Properties of subgradient projection iteration when applying to linear imaging system (Q2329650) (← links)
- A concentration bound for \(\operatorname{LSPE}( \lambda )\) (Q2677709) (← links)
- Approximate policy iteration: a survey and some new methods (Q2887629) (← links)
- Finite-Time Performance of Distributed Temporal-Difference Learning with Linear Function Approximation (Q4999359) (← links)
- A Finite Time Analysis of Temporal Difference Learning with Linear Function Approximation (Q5003727) (← links)
- Allocating resources via price management systems: a dynamic programming-based approach (Q5018825) (← links)
- Approximation of average cost Markov decision processes using empirical distributions and concentration inequalities (Q5265786) (← links)
- A Lyapunov-based version of the value iteration algorithm formulated as a discrete-time switched affine system (Q6105421) (← links)