Chapter 1 Basic Concepts of Reinforcement Learning
Chapter 1 Basic Concepts of Reinforcement Learning
Reinforcement Learning (RL) can be described by the grid world example.
We place one agent in an environment, the goal of the agent is to find a good route to the target. Every cell/grid the agent placed can be seen as a state. Agent can take one action at each state according to a certain policy. The goal of RL is to find a good policy to guide the agent taking a sequence of acitons, travelling from the start place, moving from one state to another, and finally reach the target.

Basic Concepts from the perspective of Markov decision process
Reinforcement learning utilizes the MDP framework to model the interaction between a learning agent and its environment. Actually the above example is just one simple description of the Markov decision preocess. It reflects how the agent interacts with the environment. The process directly involves three basic concepts state, action, policy and one transition state transition. It also implictly includes another concept reward.
We will introduce the key concepts one by one in the following.
State
State is the status of the agent with respect to the environment. Its set is the state space .
Action
Action is what the agent do at a certain state. The agent will obtain a new state after taking one aciton. Similarly, its set is the action space of a state denoted as .
Policy
Policy denoted as , tells the agent what actions to take at a state. It gives the probability of each action to be taken at a certain state denoted as . In mathematical form ,we use tabular representation to display one policy. In programming, we use one array, matrix to represent a policy.

Reward
Reward guides the agent toward the task objective. The agent aims to maximize expected cumulative return, whether rewards are positive or negative. For example, a return of -1 is better than -10. If the task is expressed in terms of costs instead, we minimize cumulative cost, or equivalently maximize its negative.
Reward can depend on the current state, action, and next state, for example . More generally, the joint distribution describes both the next state and reward. Writing means that the next-state and reward randomness have already been averaged out; it does not forbid rewards such as a bonus for reaching a goal.
Probability Distribution
Involve two probability form:
State transition probability: at state , taking action , the probability to transit to state is
Reward probability: at state , taking action , the probability to get reward is
Markov Property
Memoryless property: The state transiting to the next depends on current state and action rather than previous.
Other concepts
Trajectory and Return
A trajectory is a state-action-reward chain:

The return of a trajectory is the sum of its rewards. For example, a finite episode with rewards has the undiscounted return below. If the episode is extended with an absorbing terminal state, all rewards after termination are zero.
Return can be used to evaluate a policy.
Discounted Rate
The discount factor is usually for continuing tasks. Finite-horizon episodic tasks can also use .
With bounded rewards, keeps an infinite discounted sum finite. It also controls the relative weighting of near and distant rewards:
If is close to 0, the value of the discounted return is dominated by the rewards obtained in the near future.
If is close to 1, distant rewards decay more slowly. They do not necessarily dominate the return, since their magnitudes also matter.
The first reward has exponent zero:
For a continuing reward sequence , this gives
If the episode instead terminates after the first reward of 1, the discounted return is just .
Episode
When interacting with the environment following a policy, the agent may stop at some terminal states. The resulting trajectory is called an episode (or a trial), e.g. ,
An episode is usually assumed to be a finite trajectory. Tasks with episodes are called episodic tasks.
Two common ways to model a goal state should be distinguished:
Option 1: Treat the target state as a special absorbing state. Once the agent reaches an absorbing state, it will never leave. After the reward for the terminal transition, all subsequent rewards are zero. This preserves the return of the original episodic task.
Option 2: Treat the target state as a normal state with a policy. The agent can still leave the target state and gain r = +1 when entering the target state. This generally defines a different, continuing task because the agent can collect the goal reward repeatedly.
The grid-world examples in this tutorial use option 2 and treat the target as an ordinary state. It should not be confused with the zero-reward absorbing-state extension of an episodic task.
