Pages that link to "Item:Q5477859"
From MaRDI portal
The following pages link to Linear least-squares algorithms for temporal difference learning (Q5477859):
Displaying 50 items.
- Potential-based least-squares policy iteration for a parameterized feedback control system (Q289143) (← links)
- Approximate dynamic programming for the dispatch of military medical evacuation assets (Q323422) (← links)
- Perspectives of approximate dynamic programming (Q333093) (← links)
- Batch mode reinforcement learning based on the synthesis of artificial trajectories (Q378762) (← links)
- The optimal unbiased value estimator and its relation to LSTD, TD and MC (Q415609) (← links)
- Asymptotic analysis of value prediction by well-specified and misspecified models (Q448322) (← links)
- Q-learning for continuous-time linear systems: A model-free infinite horizon optimal control approach (Q511735) (← links)
- A two-level optimization model for elective surgery scheduling with downstream capacity constraints (Q666974) (← links)
- Proximal algorithms and temporal difference methods for solving fixed point problems (Q721950) (← links)
- Solving factored MDPs using non-homogeneous partitions (Q814475) (← links)
- A generalized Kalman filter for fixed point approximation and efficient temporal-difference learning (Q859737) (← links)
- Reinforcement learning algorithms with function approximation: recent advances and applications (Q903601) (← links)
- Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path (Q1009248) (← links)
- Projected equation methods for approximate solution of large linear systems (Q1012492) (← links)
- Natural actor-critic algorithms (Q1049136) (← links)
- Analytical mean squared error curves for temporal difference learning (Q1266172) (← links)
- On average versus discounted reward temporal-difference learning (Q1604814) (← links)
- Technical update: Least-squares temporal difference learning (Q1604819) (← links)
- An online prediction algorithm for reinforcement learning with linear function approximation using cross entropy method (Q1631797) (← links)
- Approximate dynamic programming for missile defense interceptor fire control (Q1751900) (← links)
- Off-policy temporal difference learning with distribution adaptation in fast mixing chains (Q1797759) (← links)
- Average cost temporal-difference learning (Q1805802) (← links)
- The convergence of \(TD(\lambda)\) for general \(\lambda\) (Q1812934) (← links)
- Least squares policy evaluation algorithms with linear function approximation (Q1870310) (← links)
- Linear least-squares algorithms for temporal difference learning (Q1911340) (← links)
- On the worst-case analysis of temporal-difference learning algorithms (Q1911342) (← links)
- Hybrid least-squares algorithms for approximate policy evaluation (Q1959511) (← links)
- Multikernel recursive least-squares temporal difference learning (Q1990335) (← links)
- Concentration bounds for temporal difference learning with linear function approximation: the case of batch data and uniform sampling (Q2051259) (← links)
- Challenges of real-world reinforcement learning: definitions, benchmarks and analysis (Q2071388) (← links)
- Improving defensive air battle management by solving a stochastic dynamic assignment problem via approximate dynamic programming (Q2103047) (← links)
- Approximate dynamic programming for the military inventory routing problem (Q2173135) (← links)
- A Q-learning predictive control scheme with guaranteed stability (Q2220029) (← links)
- An approximate dynamic programming approach for comparing firing policies in a networked air defense environment (Q2297577) (← links)
- Reinforcement learning for a biped robot based on a CPG-actor-critic method (Q2383520) (← links)
- Restricted gradient-descent algorithm for value-function approximation in reinforcement learning (Q2389624) (← links)
- Dynamic portfolio choice: a simulation-and-regression approach (Q2402578) (← links)
- Basis function adaptation in temporal difference reinforcement learning (Q2485935) (← links)
- Learning to select branching rules in the DPLL procedure for satisfiability (Q2741536) (← links)
- Convergence of the standard RLS method and<b><i>UDU</i></b><sup><i>T</i></sup>factorisation of covariance matrix for solving the algebraic Riccati equation of the DLQR via heuristic approximate dynamic programming (Q2792939) (← links)
- Chaotic dynamics and convergence analysis of temporal difference algorithms with bang-bang control (Q2800471) (← links)
- True online temporal-difference learning (Q2834469) (← links)
- Approximate policy iteration: a survey and some new methods (Q2887629) (← links)
- A review of stochastic algorithms with continuous value function approximation and some new approximate policy iteration algorithms for multidimensional continuous applications (Q2887630) (← links)
- Kalman Temporal Differences (Q3055813) (← links)
- A least squares temporal difference actor–critic algorithm with applications to warehouse management (Q3120552) (← links)
- Learning algorithms based on linearization (Q4211341) (← links)
- Artificial Intelligence and Soft Computing - ICAISC 2004 (Q4666254) (← links)
- 10.1162/1532443041827907 (Q4826001) (← links)
- (Q5168869) (← links)