Sciweavers

Free Online Productivity Tools i2Speak i2Symbol i2OCR iTex2Img iWeb2Print iWeb2Shot i2Type iPdf2Split iPdf2Merge i2Bopomofo i2Arabic i2Style i2Image i2PDF iLatex2Rtf Sci2ools

11

NIPS
1998

favoriteEmaildiscussreport

137views Information Technology» more NIPS 1998»

Risk Sensitive Reinforcement Learning

13 years 5 months ago

Risk Sensitive Reinforcement Learning

Download www.cs.cmu.edu

In this paper, we consider Markov Decision Processes (MDPs) with error states. Error states are those states entering which is undesirable or dangerous. We define the risk with respect to a policy as the probability of entering such a state when the policy is pursued. We consider the problem of finding good policies whose risk is smaller than some user-specified threshold, and formalize it as a constrained MDP with two criteria. The first criterion corresponds to the value function originally given. We will show that the risk can be formulated as a second criterion function based on a cumulative return, whose definition is independent of the original value function. We present a model free, heuristic reinforcement learning algorithm that aims at finding good deterministic policies. It is based on weighting the original value function and the risk. The weight parameter is adapted in order to find a feasible solution for the constrained problem that has a good performance with respect t...

Ralph Neuneier, Oliver Mihatsch

Real-time Traffic

Error States | NIPS 1998 | NIPS 2007 | Original Value Function | Value Function |

claim paper

Related Content

» Reinforcement learning with limited reinforcement using Bayes risk for active learning in ...

» Evolution of Reinforcement Learning in Uncertain Environments Emergence of RiskAversion an...

» RiskSensitive Online Learning

» Sensitive Discount Optimality Unifying Discounted and Average Reward Reinforcement Learnin...

» Reinforcement learning for quasipassive dynamic walking of an unstable biped robot

» Efficient methods for nearoptimal sequential decision making under uncertainty

» On step sizes stochastic shortest paths and survival probabilities in Reinforcement Learni...

» Derivatives of Logarithmic Stationary Distributions for Policy Gradient Reinforcement Lear...

» Least absolute policy iteration for robust value function approximation

Post Info
More Details (n/a)

Added	01 Nov 2010
Updated	01 Nov 2010
Type	Conference
Year	1998
Where	NIPS
Authors	Ralph Neuneier, Oliver Mihatsch

Comments (0)