You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于强化学习中有限/无限时域最优策略特性与平稳策略含义的技术问询

Great questions—let's unpack these one by one, since they're both rooted in how time horizons and MDP dynamics shape optimal behavior.

1. Why optimal policies are non-stationary in finite horizons, but stationary in infinite horizons?

Let's start with finite horizons first: think of it like a game with a hard end point. Your remaining time changes the value of each action, even if you're in the exact same state.

For example, suppose you're in a state where:

  • Action A gives you an immediate reward of 10, but drops you into a dead state where you get no more rewards.
  • Action B gives you an immediate reward of 2, but leads to a state where you get 5 rewards every step afterward.

If you only have 1 step left, Action A is clearly better (10 > 2). But if you have 3 steps left, Action B gives you 2 + 5 + 5 = 12, which beats Action A's 10. The time remaining (or equivalently, the current step in the finite horizon) directly impacts your optimal choice—so your policy has to depend on time, making it non-stationary.

Now for infinite horizons: there's no "end" to the process. At any time t, the future looks exactly the same as it does at any other time k—there's no "countdown" to change how you value future rewards. For a given state, the optimal action today will still be optimal tomorrow, next week, or 100 steps from now, because the long-term reward structure never changes.

This holds true especially when we use a discount factor γ ∈ (0,1) (to ensure the total discounted reward converges). Without a discount (γ=1), we can still have stationary optimal policies under average reward criteria, but discounted infinite horizons are the most common scenario where stationary policies are optimal.

2. What does a stationary policy mean in MDPs (given stationary dynamics)?

First, let's clarify the stationary dynamics you noted:

$$P(s_{t+1}=j|s_t=i,a) = P (s_{k+1}=j|s_k=i,a) \ \forall t,k,i,j,a$$

This describes a stationary environment: the probability of moving from state i to j when taking action a never changes, no matter what time step you're in. The environment doesn't "age" or shift its rules over time.

A stationary policy builds on this: it's a policy where your choice of action depends only on the current state, not on the time step. In plain terms: if you're in state "thirsty" on day 1, you choose "drink water"; if you're in "thirsty" on day 100, you still choose "drink water". The action for a state never varies with time.

To contrast, a non-stationary policy would let you pick different actions for the same state at different times. For example, maybe on day 1 you choose "drink water" when thirsty, but on day 10 you choose "wait for lunch" because you know lunch is coming soon (a time-dependent consideration).

Crucially: stationary environment dynamics don't force you to use a stationary policy—you could use a non-stationary one if you wanted. But in infinite horizon MDPs, the optimal policy will always be stationary, since time doesn't affect the long-term value of actions. In finite horizons, even with a stationary environment, optimal policies are non-stationary because the remaining time changes the tradeoff between immediate and future rewards.


内容的提问来源于stack exchange,提问作者niko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:44:05