跳至主要內容

Chapter 10 Actor-Critic Methods

RyanLee_ljx...大约 12 分钟RL

Chapter 10 Actor-Critic Methods

In chapter 8, we introduce value approximation function, that is to replace tabular representations for state/action value with function. Similarly, in chapter 9 we use function to represent policy instead fo tabular and turn to policy-based methods. So in this chapter we combine both of them, representing both value and policy with function and incorporating both policy-based and value-based methods.

Here, an “actor” refers to a policy update step. The reason that it is called an actor is that the actions are taken by following the policy. Here, an “critic” refers to a value update step. It is called a critic because it criticizes the actor by evaluating its corresponding values. From another point of view, actor-critic methods are still policy gradient algorithms. They can be obtained by extending the policy gradient algorithm introduced in Chapter 9.

Q actor-critic (QAC)

Actor-critic still optimizes a policy objective. For the equations in this section, use Chapter 9's discounted convention: J(θ)=(1−γ)d0TvπJ(\theta)=(1-\gamma)d_0^Tv_\pi, with fixed initial distribution d0d_0, and let η\eta be its normalized discounted state occupancy. This makes the following expectation an exact policy-gradient expression under the same assumptions. A different objective or state distribution requires corresponding changes.

θt+1=θt+α∇θJ(θt)=θt+αES∼η,A∼π[∇θln⁡π(A∣S,θt)qπ(S,A)](1) \begin{aligned} \theta_{t+1} &= \theta_t + \alpha \nabla_\theta J(\theta_t) \\ &= \theta_t + \alpha \mathbb{E}_{S \sim \eta, A \sim \pi} \left[ \nabla_\theta \ln \pi(A|S, \theta_t) q_\pi(S, A) \right] \end{aligned} \tag{1}

The stochastic gradient-ascent algorithm is

θt+1=θt+α∇θln⁡π(at∣st,θt)qt(st,at).(2) \theta_{t+1} = \theta_t + \alpha \nabla_\theta \ln \pi(a_t|s_t, \theta_t) \color{blue}{q_t(s_t, a_t)}. \tag{2}

  • If qt(st,at)q_t(s_t, a_t) is estimated by Monte Carlo learning, the corresponding algorithm is called REINFORCE or Monte Carlo policy gradient, which has already been introduced in Chapter 9 (but without value approximation function).

  • If qt(st,at)q_t(s_t, a_t) is estimated by TD learning, the corresponding algorithms are usually called actor-critic. Therefore, actor-critic methods can be obtained by incorporating TD-based value estimation into policy gradient methods.

QAC
QAC

The critic corresponds to the value update step via the Sarsa algorithm. The action values are represented by a parameterized function q(s,a,w)q(s,a,w). The actor corresponds to the policy update step in (2).

Advantage actor-critic (A2C)

The core idea of this algorithm is to introduce a baseline to reduce estimation variance.

How to do this? We need first learn about the property of baseline:

ES∼η,A∼π[∇θln⁡π(A∣S,θt)qπ(S,A)]=ES∼η,A∼π[∇θln⁡π(A∣S,θt)(qπ(S,A)−b(S))],(3) \mathbb{E}_{S \sim \eta, A \sim \pi} \left[ \nabla_\theta \ln \pi(A|S, \theta_t) q_\pi(S, A) \right] = \mathbb{E}_{S \sim \eta, A \sim \pi} \left[ \nabla_\theta \ln \pi(A|S, \theta_t) (q_\pi(S, A) - b(S)) \right], \tag{3}

where the additional baseline b(S)b(S) is a scalar function of SS. If the equation holds true, we only need to prove

ES∼η,A∼π[∇θln⁡π(A∣S,θt)b(S)]=0. \mathbb{E}_{S \sim \eta, A \sim \pi} \left[ \nabla_\theta \ln \pi(A|S, \theta_t) b(S) \right] = 0.

This equation is valid because

ES∼η,A∼π[∇θln⁡π(A∣S,θt)b(S)]=∑s∈Sη(s)∑a∈Aπ(a∣s,θt)∇θln⁡π(a∣s,θt)b(s)=∑s∈Sη(s)∑a∈A∇θπ(a∣s,θt)b(s)=∑s∈Sη(s)b(s)∑a∈A∇θπ(a∣s,θt)=∑s∈Sη(s)b(s)∇θ∑a∈Aπ(a∣s,θt)=∑s∈Sη(s)b(s)∇θ1=0. \begin{aligned} \mathbb{E}_{S \sim \eta, A \sim \pi} \left[ \nabla_\theta \ln \pi(A|S, \theta_t) b(S) \right] &= \sum_{s \in \mathcal{S}} \eta(s) \sum_{a \in \mathcal{A}} \pi(a|s, \theta_t) \nabla_\theta \ln \pi(a|s, \theta_t) b(s) \\ &= \sum_{s \in \mathcal{S}} \eta(s) \sum_{a \in \mathcal{A}} \nabla_\theta \pi(a|s, \theta_t) b(s) \\ &= \sum_{s \in \mathcal{S}} \eta(s) b(s) \sum_{a \in \mathcal{A}} \nabla_\theta \pi(a|s, \theta_t) \\ &= \sum_{s \in \mathcal{S}} \eta(s) b(s) \nabla_\theta \sum_{a \in \mathcal{A}} \pi(a|s, \theta_t) \\ &= \sum_{s \in \mathcal{S}} \eta(s) b(s) \nabla_\theta 1 = 0. \end{aligned}

Then such a expression has the same expectation with our previous metrics. Our next step is to choose a proper b(s)b(s) to reduce the approximation variance when we use samples to approximate the true gradient. Let:

X(S,A)≐∇θln⁡π(A∣S,θt)[qπ(S,A)−b(S)].(4) X(S, A) \doteq \nabla_\theta \ln \pi(A|S, \theta_t)[q_\pi(S, A) - b(S)]. \tag{4}

Then, the true gradient is E[X(S,A)]\mathbb{E}[X(S, A)]. Since we need to use a stochastic sample xx to approximate E[X]\mathbb{E}[X], it would be favorable if the variance var(X)\text{var}(X) is small. For example, if var(X)\text{var}(X) is close to zero, then any sample xx can accurately approximate E[X]\mathbb{E}[X]. On the contrary, if var(X)\text{var}(X) is large, the value of a sample may be far from E[X]\mathbb{E}[X].Although E[X]\mathbb{E}[X] is invariant to the baseline, the variance var(X)\text{var}(X) is not. Our goal is to design a good baseline to minimize var(X)\text{var}(X). In the algorithms of REINFORCE and QAC, we set b=0b = 0, which is not guaranteed to be a good baseline.

In fact, the optimal baseline that minimizes is:

b∗(s)=EA∼π[∥∇θln⁡π(A∣s,θt)∥2qπ(s,A)]EA∼π[∥∇θln⁡π(A∣s,θt)∥2],s∈S.(5) b^*(s) = \frac{\mathbb{E}_{A \sim \pi} \left[ \|\nabla_\theta \ln \pi(A|s, \theta_t)\|^2 q_\pi(s, A) \right]}{\mathbb{E}_{A \sim \pi} \left[ \|\nabla_\theta \ln \pi(A|s, \theta_t)\|^2 \right]}, \quad s \in \mathcal{S}. \tag{5}

Although the baseline in (5) is optimal, it is too complex to be useful in practice. If the weight ∥∇θln⁡π(A∣s,θt)∥2\|\nabla_\theta \ln \pi(A|s, \theta_t)\|^2 is removed from (5), we can obtain a suboptimal baseline that has a concise expression:

b†(s)=EA∼π[qπ(s,A)]=vπ(s),s∈S. b^\dagger(s) = \mathbb{E}_{A \sim \pi}[q_\pi(s, A)] = v_\pi(s), \quad s \in \mathcal{S}.

This suboptimal baseline is indeed the state value!

When b(s)=vπ(s)b(s) = v_\pi(s), the gradient-ascent algorithm in (1) becomes

θt+1=θt+αE[∇θln⁡π(A∣S,θt)[qπ(S,A)−vπ(S)]]≐θt+αE[∇θln⁡π(A∣S,θt)δπ(S,A)].(7) \begin{aligned} \theta_{t+1} &= \theta_t + \alpha \mathbb{E} \left[ \nabla_\theta \ln \pi(A|S, \theta_t)[q_\pi(S, A) - v_\pi(S)] \right] \\ &\doteq \theta_t + \alpha \mathbb{E} \left[ \nabla_\theta \ln \pi(A|S, \theta_t)\delta_\pi(S, A) \right]. \end{aligned} \tag{7}

Here,

δπ(S,A)≐qπ(S,A)−vπ(S) \delta_\pi(S, A) \doteq q_\pi(S, A) - v_\pi(S)

is called the advantage function, which reflects the advantage of one action over the others. More specifically, note that vπ(s)=∑a∈Aπ(a∣s)qπ(s,a)v_\pi(s) = \sum_{a \in \mathcal{A}} \pi(a|s) q_\pi(s, a) is the mean of the action values. If δπ(s,a)>0\delta_\pi(s, a) > 0, it means that the corresponding action has a greater value than the mean value. The stochastic version of (7) is

θt+1=θt+α∇θln⁡π(at∣st,θt)[qt(st,at)−vt(st)]=θt+α∇θln⁡π(at∣st,θt)δt(st,at),(8) \begin{aligned} \theta_{t+1} &= \theta_t + \alpha \nabla_\theta \ln \pi(a_t|s_t, \theta_t)[q_t(s_t, a_t) - v_t(s_t)] \\ &= \theta_t + \alpha \nabla_\theta \ln \pi(a_t|s_t, \theta_t)\delta_t(s_t, a_t), \end{aligned} \tag{8}

where st,ats_t, a_t are samples of S,AS, A at time tt. Here, qt(st,at)q_t(s_t, a_t) and vt(st)v_t(s_t) are approximations of qπ(θt)(st,at)q_{\pi(\theta_t)}(s_t, a_t) and vπ(θt)(st)v_{\pi(\theta_t)}(s_t), respectively.

The algorithm in (8) updates the policy based on the relative value of qtq_t with respect to vtv_t rather than the absolute value of qtq_t. This is intuitively reasonable because, when we attempt to select an action at a state, we only care about which action has the greatest value relative to the others. This can be further interpreted by:

θt+1=θt+α∇θln⁡π(at∣st,θt)δt(st,at)=θt+α∇θπ(at∣st,θt)π(at∣st,θt)δt(st,at)=θt+α(δt(st,at)π(at∣st,θt))⏟step size∇θπ(at∣st,θt) \begin{aligned} \theta_{t+1} &= \theta_t + \alpha \nabla_\theta \ln \pi(a_t|s_t, \theta_t) \delta_t(s_t, a_t) \\ &= \theta_t + \alpha \frac{\nabla_\theta \pi(a_t|s_t, \theta_t)}{\pi(a_t|s_t, \theta_t)} \delta_t(s_t, a_t) \\ &= \theta_t + \alpha \underbrace{\left( \frac{\color{blue}{\delta_t(s_t, a_t)}}{\pi(a_t|s_t, \theta_t)} \right)}_{\text{step size}} \nabla_\theta \pi(a_t|s_t, \theta_t) \end{aligned}

The update uses the relative value δt\delta_t to evaluate the sampled action against the current policy's average. This does not automatically guarantee sufficient exploration. In particular, the denominator π(at∣st,θt)\pi(a_t|s_t,\theta_t) must be considered together with ∇θπ\nabla_\theta\pi; for softmax logits, the log-probability gradient is bounded, as discussed in Chapter 9.

  • If qt(st,at)q_t(s_t, a_t) and vt(st)v_t(s_t) are estimated by Monte Carlo learning, the algorithm in (10.8) is called REINFORCE with a baseline.

  • If qt(st,at)q_t(s_t, a_t) and vt(st)v_t(s_t) are estimated by TD learning, the algorithm is usually called advantage actor-critic (A2C).

In a one-step implementation, a sampled TD error estimates the advantage:

A^t=δt=rt+1+γ(1−dt)vt(st+1)−vt(st), \hat A_t=\delta_t=r_{t+1}+\gamma(1-d_t)v_t(s_{t+1})-v_t(s_t),

where dtd_t indicates true termination. A single TD error is not the exact advantage. With the true value function and the terminal-state value set to zero, its conditional expectation is correct:

The Bellman relation is:

qπ(st,at)=E[Rt+1+γvπ(St+1)∣St=st,At=at], q_\pi(s_t, a_t) = \mathbb{E} \left[ R_{t+1} + \gamma v_\pi(S_{t+1})| S_t = s_t, A_t = a_t \right],

Thus:

qπ(st,at)−vπ(st)=E[Rt+1+γvπ(St+1)−vπ(St)∣St=st,At=at], q_\pi(s_t, a_t) - v_\pi(s_t) = \mathbb{E} \left[ R_{t+1} + \gamma v_\pi(S_{t+1}) - v_\pi(S_t) | S_t = s_t, A_t = a_t \right],

With an approximate critic, the advantage estimate can be biased. A value-only critic avoids separately estimating both Q and V: the observed reward and successor state supply a sample of the Q backup. The actor and critic can be separate networks or share a feature backbone with different heads.

For one transition, let yt=rt+1+γ(1−dt)v(st+1,w)y_t=r_{t+1}+\gamma(1-d_t)v(s_{t+1},w). The corresponding losses are:

Lactor=−log⁡π(at∣st,θ)stopgrad⁡(yt−v(st,w)),Lcritic=12(v(st,w)−stopgrad⁡(yt))2. L_{actor}=-\log\pi(a_t|s_t,\theta)\operatorname{stopgrad}(y_t-v(s_t,w)),\qquad L_{critic}=\frac12\big(v(s_t,w)-\operatorname{stopgrad}(y_t)\big)^2.

The actor treats the advantage as a fixed evaluation signal; it must not reduce its loss by changing the critic's score through this term. The critic stops gradients through its bootstrap target, giving the semi-gradient update from Chapter 8. Minibatch and time weighting must match the chosen policy objective and sampling convention. Practical A2C commonly collects synchronous rollouts from multiple environments and uses n-step returns or other advantage estimates.

This one-step method is also called TD actor-critic. Its stochastic actor can generate actions directly without an additional ϵ\epsilon-greedy rule, although useful exploration and sufficient state coverage are still not guaranteed. Entropy regularization is often used to discourage premature concentration of the policy.

A2C
A2C

Off-policy actor-critic

Importance sampling

Consider a random variable X∈XX \in \mathcal{X}. Suppose that p0(X)p_0(X) is a probability distribution. Our goal is to estimate EX∼p0[X]\mathbb{E}_{X \sim p_0}[X]. Suppose that we have some i.i.d. samples {xi}i=1n\{x_i\}_{i=1}^n.

  • First, if the samples {xi}i=1n\{x_i\}_{i=1}^n are generated by following p0p_0, then the average value xˉ=1n∑i=1nxi\bar{x} = \frac{1}{n} \sum_{i=1}^n x_i can be used to approximate EX∼p0[X]\mathbb{E}_{X \sim p_0}[X] because xˉ\bar{x} is an unbiased estimate of EX∼p0[X]\mathbb{E}_{X \sim p_0}[X] and the estimation variance converges to zero as n→∞n \to \infty (see the law of large numbers in Box 5.1 for more information).

  • Second, consider a new scenario where the samples {xi}i=1n\{x_i\}_{i=1}^n are not generated by p0p_0. Instead, they are generated by another distribution p1p_1. Can we still use these samples to approximate EX∼p0[X]\mathbb{E}_{X \sim p_0}[X]? The answer is yes. However, we can no longer use xˉ=1n∑i=1nxi\bar{x} = \frac{1}{n} \sum_{i=1}^n x_i to approximate EX∼p0[X]\mathbb{E}_{X \sim p_0}[X] since xˉ≈EX∼p1[X]\bar{x} \approx \mathbb{E}_{X \sim p_1}[X] rather than EX∼p0[X]\mathbb{E}_{X \sim p_0}[X].

In the second scenario, importance sampling can be used provided p1(x)>0p_1(x)>0 wherever p0(x)>0p_0(x)>0, with the required integrability conditions. In particular:

EX∼p0[X]=∑x∈Xp0(x)x=∑x∈Xp1(x)p0(x)p1(x)x⏟f(x)=EX∼p1[f(X)].(9) \mathbb{E}_{X \sim p_0}[X] = \sum_{x \in \mathcal{X}} p_0(x)x = \sum_{x \in \mathcal{X}} p_1(x) \underbrace{\frac{p_0(x)}{p_1(x)} x}_{f(x)} = \mathbb{E}_{X \sim p_1}[f(X)]. \tag{9}

Thus, estimating EX∼p0[X]\mathbb{E}_{X \sim p_0}[X] becomes the problem of estimating EX∼p1[f(X)]\mathbb{E}_{X \sim p_1}[f(X)]. Let:

fˉ≐1n∑i=1nf(xi). \bar{f} \doteq \frac{1}{n} \sum_{i=1}^n f(x_i).

Since fˉ\bar{f} can effectively approximate EX∼p1[f(X)]\mathbb{E}_{X \sim p_1}[f(X)], it then follows from (9) that

EX∼p0[X]=EX∼p1[f(X)]≈fˉ=1n∑i=1nf(xi)=1n∑i=1np0(xi)p1(xi)⏟importance weightxi.(10) \mathbb{E}_{X \sim p_0}[X] = \mathbb{E}_{X \sim p_1}[f(X)] \approx \bar{f} = \frac{1}{n} \sum_{i=1}^n f(x_i) = \frac{1}{n} \sum_{i=1}^n \underbrace{\frac{p_0(x_i)}{p_1(x_i)}}_{\text{importance weight}} x_i. \tag{10}

Equation (10) suggests that EX∼p0[X]\mathbb{E}_{X \sim p_0}[X] can be approximated by a weighted average of xix_i. Here, p0(xi)p1(xi)\frac{p_0(x_i)}{p_1(x_i)} is called the importance weight. Why is it called importance sampling? Because we want to find p0p_0. If we sample a xix_i, and its probability is high under p0p_0 but low under p1p_1, it means that it appears a lot under p0p_0 but very little under the current p1p_1 sampling. Therefore, we should cherish this hard-won sample, that is, this sample is very important, and we give it a large weight.

A common RL reason for using p1p_1 samples is that available data were generated by an older or different behavior policy. A neural-network target policy can often be sampled directly; its representation is not itself a reason to require importance sampling. Large importance weights can make the estimator's variance high, and finite-variance claims require appropriate moment assumptions.

You can see the illustrative example provided in the tutorial to better understand it.

The off-policy policy gradient theorem

Like the previous on-policy case, we need to derive the policy gradient in the off-policy case.

Suppose β\beta is the behavior policy that generates experience samples. Our aim is to use these samples to update a target policy π\pi that can maximize the metric:

J(θ)=∑s∈Sdβ(s)vπ(s)=ES∼dβ[vπ(S)], J(\theta) = \sum_{s \in \mathcal{S}} d_\beta(s) v_\pi(s) = \mathbb{E}_{S \sim d_\beta}[v_\pi(S)],

Here dβd_\beta is a stationary distribution under a fixed behavior policy β\beta and is held independent of θ\theta when differentiating. Unlike the normalized objective in the on-policy section, this J is written without the factor 1−γ1-\gamma.

Theorem 10.1 (Off-policy policy gradient theorem)

In the discounted case where γ∈(0,1)\gamma \in (0, 1), the gradient of J(θ)J(\theta) is

∇θJ(θ)=11−γES∼ρˉ,A∼β[π(A∣S,θ)β(A∣S)⏟action importance weight∇θln⁡π(A∣S,θ)qπ(S,A)],(11) \nabla_\theta J(\theta) = \frac{1}{1-\gamma}\mathbb{E}_{S \sim \bar\rho, A \sim \beta} \left[ \underbrace{\frac{\pi(A|S, \theta)}{\beta(A|S)}}_{\text{action importance weight}} \nabla_\theta \ln \pi(A|S, \theta) q_\pi(S, A) \right], \tag{11}

where ρ\rho is an unnormalized discounted occupancy weight:

ρ(s)≐∑s′∈Sdβ(s′)Prπ(s∣s′),s∈S, \rho(s) \doteq \sum_{s' \in \mathcal{S}} d_\beta(s') \text{Pr}_\pi(s|s'), \quad s \in \mathcal{S},

where Prπ(s∣s′)=∑k=0∞γk[Pπk]s′s=[(I−γPπ)−1]s′s\text{Pr}_\pi(s|s') = \sum_{k=0}^\infty \gamma^k [P_\pi^k]_{s's} = [(I - \gamma P_\pi)^{-1}]_{s's} is the discounted total probability of transitioning from s′s' to ss under policy π\pi.

Its total mass is ∑sρ(s)=1/(1−γ)\sum_s\rho(s)=1/(1-\gamma), so ρˉ(s)=(1−γ)ρ(s)\bar\rho(s)=(1-\gamma)\rho(s) is the normalized state distribution used in the expectation. The action ratio corrects the action distribution at a given state; it does not replace the required state distribution ρˉ\bar\rho with the behavior distribution dβd_\beta. If states are instead sampled from dβd_\beta, an exact expression for this objective also needs the state-density ratio ρˉ(s)/dβ(s)\bar\rho(s)/d_\beta(s) wherever it is required, together with state and action coverage. In episodic settings, zero-reward absorbing states give the same normalization convention.

The algorithm of off-policy actor-critic

Subtracting an action-independent baseline gives the following action-weighted update form, with the constant 1/(1−γ)1/(1-\gamma) absorbed into the learning rate. It corresponds to the theorem's gradient only when the required state weighting and value estimates are also handled correctly:

θt+1=θt+αθπ(at∣st,θt)β(at∣st)∇θln⁡π(at∣st,θt)δt(st,at) \theta_{t+1} = \theta_t + \alpha_\theta \frac{\pi(a_t|s_t, \theta_t)}{\beta(a_t|s_t)} \nabla_\theta \ln \pi(a_t|s_t, \theta_t) \delta_t(s_t, a_t)

and hence:

θt+1=θt+αθ(δt(st,at)β(at∣st))∇θπ(at∣st,θt) \theta_{t+1} = \theta_t + \alpha_\theta \left( \frac{\delta_t(s_t, a_t)}{\beta(a_t|s_t)} \right) \nabla_\theta \pi(a_t|s_t, \theta_t)

When the data are ordinary behavior-policy trajectories, the update above alone is generally an approximation to the stated objective's gradient: it omits the state-distribution correction. It should not be interpreted as a universally exact off-policy gradient merely because an action ratio has been added. The critic also needs an appropriate off-policy evaluation procedure for the target policy. Read the following textbook algorithm with these qualifications.

Off-policy actor-critic based on importance sampling
Off-policy actor-critic based on importance sampling

Deterministic actor-critic (DPG)

The score-function methods above use stochastic policies with well-defined log probabilities or log densities on their sampling support. For discrete actions, softmax is a common parameterization. Continuous actions can also use stochastic policies, such as a∼N(μθ(s),Σθ(s))a\sim\mathcal N(\mu_\theta(s),\Sigma_\theta(s)); continuous-control PPO and SAC commonly do so. A deterministic policy is a separate choice, used by methods such as DDPG and TD3, rather than a consequence of having a continuous action space.

The deterministic policy is specifically denoted as:

a=μ(s,θ)≐μ(s) a = \mu(s, \theta) \doteq \mu(s)

μ\mu is a mapping from S\mathcal{S} to A\mathcal{A}. μ\mu can be represented by, for example, a neural network with the input as ss, the output as aa, and the parameter as θ\theta.We may write μ(s,θ)\mu(s, \theta) in short as μ(s)\mu(s).

The policy gradient theorems introduced before are merely valid for stochastic policies. If the policy must be deterministic, we must derive a new policy gradient theorem.

Deterministic policy gradient theorem

Returning to the normalized discounted fixed-start objective J(θ)=(1−γ)d0TvμJ(\theta)=(1-\gamma)d_0^Tv_\mu, and under the differentiability and regularity assumptions of the deterministic policy gradient theorem, the gradient is:

∇θJ(θ)=∑s∈Sη(s)∇θμ(s)(∇aqμ(s,a))∣a=μ(s)=ES∼η[∇θμ(S)(∇aqμ(S,a))∣a=μ(S)],(11) \begin{aligned} \nabla_\theta J(\theta) &= \sum_{s \in \mathcal{S}} \eta(s) \nabla_\theta \mu(s) (\nabla_a q_\mu(s, a))|_{a=\mu(s)} \\ &= \mathbb{E}_{S \sim \eta} \left[ \nabla_\theta \mu(S) (\nabla_a q_\mu(S, a))|_{a=\mu(S)} \right], \end{aligned} \tag{11}

Here η\eta is the normalized discounted occupancy induced by μ\mu from d0d_0. For an unnormalized objective, the corresponding factor 1/(1−γ)1/(1-\gamma) must be restored. State weighting is part of the theorem, not an arbitrary replay distribution.

The actor gradient evaluates the critic's action derivative at a=μ(s)a=\mu(s), so it does not need to sample an action for that gradient calculation. This enables useful off-policy implementations, but the definition of off-policy remains a mismatch between behavior and target policies. Omitting an action random variable alone does not establish that mismatch or correct a different state distribution.

Based on the gradient given in Theorem, we can apply the gradient-ascent algorithm to maximize J(θ)J(\theta):

θt+1=θt+αθES∼η[∇θμ(S)(∇aqμ(S,a))∣a=μ(S)]. \theta_{t+1} = \theta_t + \alpha_\theta \mathbb{E}_{S \sim \eta} \left[ \nabla_\theta \mu(S) (\nabla_a q_\mu(S, a))|_{a=\mu(S)} \right].

The corresponding stochastic gradient-ascent algorithm is

θt+1=θt+αθ∇θμ(st)(∇aqμ(st,a))∣a=μ(st). \theta_{t+1} = \theta_t + \alpha_\theta \nabla_\theta \mu(s_t) (\nabla_a q_\mu(s_t, a))|_{a=\mu(s_t)}.

In a common off-policy implementation, an exploratory behavior policy β\beta collects data while the target actor is μ\mu. The actor can use replay states and evaluate its current action μ(s)\mu(s), but this state weighting is not automatically the exact occupancy weighting in the fixed-start theorem above.

The critic also evaluates the target policy using behavior-policy transitions. Given a state-action pair, the conditional environment transition and reward law is the same regardless of which policy selected that action. A Q backup can therefore use the observed transition and the target actor at the next state without an action importance ratio; adequate coverage and stable function approximation are still required. In particular, the experience sample required by the critic is (st,at,rt+1,st+1,a~t+1)(s_t, a_t, r_{t+1}, s_{t+1}, \tilde{a}_{t+1}), where a~t+1=μ(st+1)\tilde{a}_{t+1} = \mu(s_{t+1}). The generation of this experience sample involves two policies. The first is the policy for generating ata_t at sts_t, and the second is the policy for generating a~t+1\tilde{a}_{t+1} at st+1s_{t+1}. The first policy that generates ata_t is the behavior policy since ata_t is used to interact with the environment. The second policy must be μ\mu because it is the policy that the critic aims to evaluate. Hence, μ\mu is the target policy. It should be noted that a~t+1\tilde{a}_{t+1} is not used to interact with the environment in the next time step. Hence, μ\mu is not the behavior policy. Therefore, the critic is off-policy.

How to select the function q(s,a,w)q(s, a, w)? The original research work [74] that proposed the deterministic policy gradient method adopted linear functions: q(s,a,w)=ϕT(s,a)wq(s, a, w) = \phi^T(s, a)w where ϕ(s,a)\phi(s, a) is the feature vector. It is currently popular to represent q(s,a,w)q(s, a, w) using neural networks, as suggested in the deep deterministic policy gradient (DDPG) method.

DPG
DPG
评论
  • 按正序
  • 按倒序
  • 按热度
Powered by Waline v3.1.3