Chapter 10 Actor-Critic Methods
Chapter 10 Actor-Critic Methods
In chapter 8, we introduce value approximation function, that is to replace tabular representations for state/action value with function. Similarly, in chapter 9 we use function to represent policy instead fo tabular and turn to policy-based methods. So in this chapter we combine both of them, representing both value and policy with function and incorporating both policy-based and value-based methods.
Here, an “actor” refers to a policy update step. The reason that it is called an actor is that the actions are taken by following the policy. Here, an “critic” refers to a value update step. It is called a critic because it criticizes the actor by evaluating its corresponding values. From another point of view, actor-critic methods are still policy gradient algorithms. They can be obtained by extending the policy gradient algorithm introduced in Chapter 9.
Q actor-critic (QAC)
Actor-critic still optimizes a policy objective. For the equations in this section, use Chapter 9's discounted convention: , with fixed initial distribution , and let be its normalized discounted state occupancy. This makes the following expectation an exact policy-gradient expression under the same assumptions. A different objective or state distribution requires corresponding changes.
The stochastic gradient-ascent algorithm is
If is estimated by Monte Carlo learning, the corresponding algorithm is called REINFORCE or Monte Carlo policy gradient, which has already been introduced in Chapter 9 (but without value approximation function).
If is estimated by TD learning, the corresponding algorithms are usually called actor-critic. Therefore, actor-critic methods can be obtained by incorporating TD-based value estimation into policy gradient methods.

The critic corresponds to the value update step via the Sarsa algorithm. The action values are represented by a parameterized function . The actor corresponds to the policy update step in (2).
Advantage actor-critic (A2C)
The core idea of this algorithm is to introduce a baseline to reduce estimation variance.
How to do this? We need first learn about the property of baseline:
where the additional baseline is a scalar function of . If the equation holds true, we only need to prove
This equation is valid because
Then such a expression has the same expectation with our previous metrics. Our next step is to choose a proper to reduce the approximation variance when we use samples to approximate the true gradient. Let:
Then, the true gradient is . Since we need to use a stochastic sample to approximate , it would be favorable if the variance is small. For example, if is close to zero, then any sample can accurately approximate . On the contrary, if is large, the value of a sample may be far from .Although is invariant to the baseline, the variance is not. Our goal is to design a good baseline to minimize . In the algorithms of REINFORCE and QAC, we set , which is not guaranteed to be a good baseline.
In fact, the optimal baseline that minimizes is:
Although the baseline in (5) is optimal, it is too complex to be useful in practice. If the weight is removed from (5), we can obtain a suboptimal baseline that has a concise expression:
This suboptimal baseline is indeed the state value!
When , the gradient-ascent algorithm in (1) becomes
Here,
is called the advantage function, which reflects the advantage of one action over the others. More specifically, note that is the mean of the action values. If , it means that the corresponding action has a greater value than the mean value. The stochastic version of (7) is
where are samples of at time . Here, and are approximations of and , respectively.
The algorithm in (8) updates the policy based on the relative value of with respect to rather than the absolute value of . This is intuitively reasonable because, when we attempt to select an action at a state, we only care about which action has the greatest value relative to the others. This can be further interpreted by:
The update uses the relative value to evaluate the sampled action against the current policy's average. This does not automatically guarantee sufficient exploration. In particular, the denominator must be considered together with ; for softmax logits, the log-probability gradient is bounded, as discussed in Chapter 9.
If and are estimated by Monte Carlo learning, the algorithm in (10.8) is called REINFORCE with a baseline.
If and are estimated by TD learning, the algorithm is usually called advantage actor-critic (A2C).
In a one-step implementation, a sampled TD error estimates the advantage:
where indicates true termination. A single TD error is not the exact advantage. With the true value function and the terminal-state value set to zero, its conditional expectation is correct:
The Bellman relation is:
Thus:
With an approximate critic, the advantage estimate can be biased. A value-only critic avoids separately estimating both Q and V: the observed reward and successor state supply a sample of the Q backup. The actor and critic can be separate networks or share a feature backbone with different heads.
For one transition, let . The corresponding losses are:
The actor treats the advantage as a fixed evaluation signal; it must not reduce its loss by changing the critic's score through this term. The critic stops gradients through its bootstrap target, giving the semi-gradient update from Chapter 8. Minibatch and time weighting must match the chosen policy objective and sampling convention. Practical A2C commonly collects synchronous rollouts from multiple environments and uses n-step returns or other advantage estimates.
This one-step method is also called TD actor-critic. Its stochastic actor can generate actions directly without an additional -greedy rule, although useful exploration and sufficient state coverage are still not guaranteed. Entropy regularization is often used to discourage premature concentration of the policy.

Off-policy actor-critic
Importance sampling
Consider a random variable . Suppose that is a probability distribution. Our goal is to estimate . Suppose that we have some i.i.d. samples .
First, if the samples are generated by following , then the average value can be used to approximate because is an unbiased estimate of and the estimation variance converges to zero as (see the law of large numbers in Box 5.1 for more information).
Second, consider a new scenario where the samples are not generated by . Instead, they are generated by another distribution . Can we still use these samples to approximate ? The answer is yes. However, we can no longer use to approximate since rather than .
In the second scenario, importance sampling can be used provided wherever , with the required integrability conditions. In particular:
Thus, estimating becomes the problem of estimating . Let:
Since can effectively approximate , it then follows from (9) that
Equation (10) suggests that can be approximated by a weighted average of . Here, is called the importance weight. Why is it called importance sampling? Because we want to find . If we sample a , and its probability is high under but low under , it means that it appears a lot under but very little under the current sampling. Therefore, we should cherish this hard-won sample, that is, this sample is very important, and we give it a large weight.
A common RL reason for using samples is that available data were generated by an older or different behavior policy. A neural-network target policy can often be sampled directly; its representation is not itself a reason to require importance sampling. Large importance weights can make the estimator's variance high, and finite-variance claims require appropriate moment assumptions.
You can see the illustrative example provided in the tutorial to better understand it.
The off-policy policy gradient theorem
Like the previous on-policy case, we need to derive the policy gradient in the off-policy case.
Suppose is the behavior policy that generates experience samples. Our aim is to use these samples to update a target policy that can maximize the metric:
Here is a stationary distribution under a fixed behavior policy and is held independent of when differentiating. Unlike the normalized objective in the on-policy section, this J is written without the factor .
Theorem 10.1 (Off-policy policy gradient theorem)
In the discounted case where , the gradient of is
where is an unnormalized discounted occupancy weight:
where is the discounted total probability of transitioning from to under policy .
Its total mass is , so is the normalized state distribution used in the expectation. The action ratio corrects the action distribution at a given state; it does not replace the required state distribution with the behavior distribution . If states are instead sampled from , an exact expression for this objective also needs the state-density ratio wherever it is required, together with state and action coverage. In episodic settings, zero-reward absorbing states give the same normalization convention.
The algorithm of off-policy actor-critic
Subtracting an action-independent baseline gives the following action-weighted update form, with the constant absorbed into the learning rate. It corresponds to the theorem's gradient only when the required state weighting and value estimates are also handled correctly:
and hence:
When the data are ordinary behavior-policy trajectories, the update above alone is generally an approximation to the stated objective's gradient: it omits the state-distribution correction. It should not be interpreted as a universally exact off-policy gradient merely because an action ratio has been added. The critic also needs an appropriate off-policy evaluation procedure for the target policy. Read the following textbook algorithm with these qualifications.

Deterministic actor-critic (DPG)
The score-function methods above use stochastic policies with well-defined log probabilities or log densities on their sampling support. For discrete actions, softmax is a common parameterization. Continuous actions can also use stochastic policies, such as ; continuous-control PPO and SAC commonly do so. A deterministic policy is a separate choice, used by methods such as DDPG and TD3, rather than a consequence of having a continuous action space.
The deterministic policy is specifically denoted as:
is a mapping from to . can be represented by, for example, a neural network with the input as , the output as , and the parameter as .We may write in short as .
The policy gradient theorems introduced before are merely valid for stochastic policies. If the policy must be deterministic, we must derive a new policy gradient theorem.
Deterministic policy gradient theorem
Returning to the normalized discounted fixed-start objective , and under the differentiability and regularity assumptions of the deterministic policy gradient theorem, the gradient is:
Here is the normalized discounted occupancy induced by from . For an unnormalized objective, the corresponding factor must be restored. State weighting is part of the theorem, not an arbitrary replay distribution.
The actor gradient evaluates the critic's action derivative at , so it does not need to sample an action for that gradient calculation. This enables useful off-policy implementations, but the definition of off-policy remains a mismatch between behavior and target policies. Omitting an action random variable alone does not establish that mismatch or correct a different state distribution.
Based on the gradient given in Theorem, we can apply the gradient-ascent algorithm to maximize :
The corresponding stochastic gradient-ascent algorithm is
In a common off-policy implementation, an exploratory behavior policy collects data while the target actor is . The actor can use replay states and evaluate its current action , but this state weighting is not automatically the exact occupancy weighting in the fixed-start theorem above.
The critic also evaluates the target policy using behavior-policy transitions. Given a state-action pair, the conditional environment transition and reward law is the same regardless of which policy selected that action. A Q backup can therefore use the observed transition and the target actor at the next state without an action importance ratio; adequate coverage and stable function approximation are still required. In particular, the experience sample required by the critic is , where . The generation of this experience sample involves two policies. The first is the policy for generating at , and the second is the policy for generating at . The first policy that generates is the behavior policy since is used to interact with the environment. The second policy must be because it is the policy that the critic aims to evaluate. Hence, is the target policy. It should be noted that is not used to interact with the environment in the next time step. Hence, is not the behavior policy. Therefore, the critic is off-policy.
How to select the function ? The original research work [74] that proposed the deterministic policy gradient method adopted linear functions: where is the feature vector. It is currently popular to represent using neural networks, as suggested in the deep deterministic policy gradient (DDPG) method.

