You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何相同Q值表下CartPole的Q-Learning测试结果不一致?

问题

我正在学习机器学习,为Gymnasium环境开发强化学习算法。已在简单问题上实现Q-Learning,随后将该算法应用到CartPole这类复杂度稍高的环境中。训练时算法表现符合预期,AI能取得不错的结果,且结果会随训练轮数、学习率、epsilon值变化而升降。但训练完成后测试时,我关闭了epsilon探索概率,理论上AI应每次采取最优动作,测试结果应一致,实际却每次结果不同。我是否对Q-Learning的原理理解有误,还是这种情况本就正常?

原因分析

这种情况是正常的,并非对Q-Learning原理理解有误,核心原因包括:

  • 状态离散化的精度损失:代码将CartPole的连续状态(小车位置、速度、杆角度、角速度)离散化为有限的区间(bins),多个不同的真实连续状态会被映射到同一个离散状态。AI在同一离散状态下选择最优动作时,不同的真实初始状态(每次env.reset()的初始状态存在微小随机差异)会导致后续状态演化路径完全不同,最终奖励差异明显。
  • 环境初始状态的随机性:CartPole每次env.reset()会生成随机的初始状态,比如杆的初始角度、小车初始速度存在微小波动。这些微小差异在离散化后可能被归为同一状态,但实际演化路径会因初始条件不同而发散,导致最终结束步数不同。
  • Q-table的局限性:训练后的Q-table仅在有限离散状态下学习到局部最优动作,并非对所有可能的连续状态都能保证完美平衡杆。当遇到Q-table未充分训练的状态演化路径时,AI的动作可能无法维持平衡,导致提前结束。
优化建议

若希望测试结果更稳定,可尝试以下调整:

  • 提升离散化精度:增加state_space的区间数量,比如将小车位置、速度的区间从20提升到50,杆角度、角速度从50提升到100,减少不同连续状态映射到同一离散状态的情况。
  • 固定测试初始状态:测试时手动设置固定的初始状态,替代env.reset()的随机初始值,示例代码:
# 替换测试环节的env.reset()为固定初始状态
fixed_initial_state = np.array([0.0, 0.0, 0.01, 0.0])
state = discretize_state(fixed_initial_state)
  • 改用函数近似替代Q-table:对于连续状态空间,使用神经网络(如DQN)近似Q值,避免离散化带来的精度损失,能更好地处理连续状态的细微差异。
附:实现代码及测试结果

代码实现

import gymnasium as gym
import numpy as np
import random

# Hyperparamteters
alpha = 0.05
gamma = 0.90
epsilon = 1.0
epsilon_decay = 0.995
epsilon_min = 0.1
episodes = 10000
max_steps = 200

# Initialise environment
env = gym.make("CartPole-v1")
state_space = [20, 20, 50, 50] #cart_position, cart_velocity, pole_angle, pole_angular_velocity
q_table = np.zeros(state_space + [env.action_space.n])

def discretize_state(state):
    """
    Discretize a space means to convert all continuous actions into discrete and finite actions. 
    Continuous action can be infinite or very large and will therefore be difficult to handle. 
    This function takes a state representing all the values for each dimension [cart position, cart velocity, pole angle, pole velocity]
    and returns a discretised tuple rounded to the closest integer. 
    """

    # Normalization formula = (state - min) / (max - min). Returns a value between 0 - 1
    normalised_state = (state - env.observation_space.low) / (env.observation_space.high - env.observation_space.low)

    # Scales the normalised values into the number and size of bins, then it rounds each direction into their closest integer value.
    discretized = np.round(normalised_state * (np.array(state_space) - 1)).astype(int)

    return tuple(discretized)

# Q-learning loop algorithm
print("Training started:\n-----------------------------------\n")
for episode in range(episodes):
    state = discretize_state(env.reset()[0])
    total_reward = 0

    for step in range(max_steps):
        """
        Decide wether to explore or exploit based on epsilon. with probability epsilon the 
        algorithm will explore by taking a random possible action. With probability 1 - epsilon
        the algorithm will take the best possible action based on the q-value of the previously
        explored actions. 
        As epsilon starts with a value of 1, the first action will always be random. 
        """
        if random.uniform(0, 1) < epsilon:
            action = env.action_space.sample()
        else:
            action = np.argmax(q_table[state])

        next_state, reward, done, _, _ = env.step(action)
        next_state = discretize_state(next_state)
        total_reward += reward

        # Q-Learning algorithm:  Q(s, a) <- Q(s, a) + alpha[R + gamma * max(Q(s1, a1)) - Q(s, a))]
        best_next_action = np.argmax(q_table[next_state])
        td_target = reward + gamma * q_table[next_state][best_next_action] # Temporal Difference Target -> sum of total reward and the discounted q value of best action for next state
        td_error = td_target - q_table[state][action] # Temporal Difference Error -> Difference between TDTarget and the current q-value
        q_table[state][action] = q_table[state][action] + alpha * td_error # Update current q-value based on the larning rate (alpha)

        state = next_state

        if done:
            break
        
    # Reduce epsilon by epsilon decay rate to gradually reduce exploration and favour learning on previous experiences
    epsilon = max(epsilon_min, epsilon * epsilon_decay)
    print(f"Epsiode {episode + 1}: Total reward: {total_reward}")
print("Training finished.")

# Test trained agend
print("Testing:\n----------------------------------\n")
for episode in range(10):
    state = discretize_state(env.reset()[0])
    total_reward = 0

    for step in range(max_steps):
        action = np.argmax(q_table[state])
        next_state, reward, done, _, _ = env.step(action)
        state = discretize_state(next_state)
        total_reward += reward

        if done:
            print(f"Episode {episode + 1} - Total reward: {total_reward}")
            break

env.close()

测试结果示例

Episode 1 - Total reward: 38.0
Episode 2 - Total reward: 46.0
Episode 3 - Total reward: 48.0
Episode 4 - Total reward: 62.0
Episode 5 - Total reward: 48.0
Episode 6 - Total reward: 65.0
Episode 7 - Total reward: 44.0
Episode 8 - Total reward: 69.0
Episode 9 - Total reward: 59.0
Episode 10 - Total reward: 16.0

内容的提问来源于stack exchange,提问作者Freddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 04:42:06