为何相同Q值表下CartPole的Q-Learning测试结果不一致?
问题
我正在学习机器学习,为Gymnasium环境开发强化学习算法。已在简单问题上实现Q-Learning,随后将该算法应用到CartPole这类复杂度稍高的环境中。训练时算法表现符合预期,AI能取得不错的结果,且结果会随训练轮数、学习率、epsilon值变化而升降。但训练完成后测试时,我关闭了epsilon探索概率,理论上AI应每次采取最优动作,测试结果应一致,实际却每次结果不同。我是否对Q-Learning的原理理解有误,还是这种情况本就正常?
原因分析
这种情况是正常的,并非对Q-Learning原理理解有误,核心原因包括:
- 状态离散化的精度损失:代码将CartPole的连续状态(小车位置、速度、杆角度、角速度)离散化为有限的区间(bins),多个不同的真实连续状态会被映射到同一个离散状态。AI在同一离散状态下选择最优动作时,不同的真实初始状态(每次
env.reset()的初始状态存在微小随机差异)会导致后续状态演化路径完全不同,最终奖励差异明显。 - 环境初始状态的随机性:CartPole每次
env.reset()会生成随机的初始状态,比如杆的初始角度、小车初始速度存在微小波动。这些微小差异在离散化后可能被归为同一状态,但实际演化路径会因初始条件不同而发散,导致最终结束步数不同。 - Q-table的局限性:训练后的Q-table仅在有限离散状态下学习到局部最优动作,并非对所有可能的连续状态都能保证完美平衡杆。当遇到Q-table未充分训练的状态演化路径时,AI的动作可能无法维持平衡,导致提前结束。
优化建议
若希望测试结果更稳定,可尝试以下调整:
- 提升离散化精度:增加
state_space的区间数量,比如将小车位置、速度的区间从20提升到50,杆角度、角速度从50提升到100,减少不同连续状态映射到同一离散状态的情况。 - 固定测试初始状态:测试时手动设置固定的初始状态,替代
env.reset()的随机初始值,示例代码:
# 替换测试环节的env.reset()为固定初始状态 fixed_initial_state = np.array([0.0, 0.0, 0.01, 0.0]) state = discretize_state(fixed_initial_state)
- 改用函数近似替代Q-table:对于连续状态空间,使用神经网络(如DQN)近似Q值,避免离散化带来的精度损失,能更好地处理连续状态的细微差异。
附:实现代码及测试结果
代码实现
import gymnasium as gym import numpy as np import random # Hyperparamteters alpha = 0.05 gamma = 0.90 epsilon = 1.0 epsilon_decay = 0.995 epsilon_min = 0.1 episodes = 10000 max_steps = 200 # Initialise environment env = gym.make("CartPole-v1") state_space = [20, 20, 50, 50] #cart_position, cart_velocity, pole_angle, pole_angular_velocity q_table = np.zeros(state_space + [env.action_space.n]) def discretize_state(state): """ Discretize a space means to convert all continuous actions into discrete and finite actions. Continuous action can be infinite or very large and will therefore be difficult to handle. This function takes a state representing all the values for each dimension [cart position, cart velocity, pole angle, pole velocity] and returns a discretised tuple rounded to the closest integer. """ # Normalization formula = (state - min) / (max - min). Returns a value between 0 - 1 normalised_state = (state - env.observation_space.low) / (env.observation_space.high - env.observation_space.low) # Scales the normalised values into the number and size of bins, then it rounds each direction into their closest integer value. discretized = np.round(normalised_state * (np.array(state_space) - 1)).astype(int) return tuple(discretized) # Q-learning loop algorithm print("Training started:\n-----------------------------------\n") for episode in range(episodes): state = discretize_state(env.reset()[0]) total_reward = 0 for step in range(max_steps): """ Decide wether to explore or exploit based on epsilon. with probability epsilon the algorithm will explore by taking a random possible action. With probability 1 - epsilon the algorithm will take the best possible action based on the q-value of the previously explored actions. As epsilon starts with a value of 1, the first action will always be random. """ if random.uniform(0, 1) < epsilon: action = env.action_space.sample() else: action = np.argmax(q_table[state]) next_state, reward, done, _, _ = env.step(action) next_state = discretize_state(next_state) total_reward += reward # Q-Learning algorithm: Q(s, a) <- Q(s, a) + alpha[R + gamma * max(Q(s1, a1)) - Q(s, a))] best_next_action = np.argmax(q_table[next_state]) td_target = reward + gamma * q_table[next_state][best_next_action] # Temporal Difference Target -> sum of total reward and the discounted q value of best action for next state td_error = td_target - q_table[state][action] # Temporal Difference Error -> Difference between TDTarget and the current q-value q_table[state][action] = q_table[state][action] + alpha * td_error # Update current q-value based on the larning rate (alpha) state = next_state if done: break # Reduce epsilon by epsilon decay rate to gradually reduce exploration and favour learning on previous experiences epsilon = max(epsilon_min, epsilon * epsilon_decay) print(f"Epsiode {episode + 1}: Total reward: {total_reward}") print("Training finished.") # Test trained agend print("Testing:\n----------------------------------\n") for episode in range(10): state = discretize_state(env.reset()[0]) total_reward = 0 for step in range(max_steps): action = np.argmax(q_table[state]) next_state, reward, done, _, _ = env.step(action) state = discretize_state(next_state) total_reward += reward if done: print(f"Episode {episode + 1} - Total reward: {total_reward}") break env.close()
测试结果示例
Episode 1 - Total reward: 38.0 Episode 2 - Total reward: 46.0 Episode 3 - Total reward: 48.0 Episode 4 - Total reward: 62.0 Episode 5 - Total reward: 48.0 Episode 6 - Total reward: 65.0 Episode 7 - Total reward: 44.0 Episode 8 - Total reward: 69.0 Episode 9 - Total reward: 59.0 Episode 10 - Total reward: 16.0
内容的提问来源于stack exchange,提问作者Freddy
相关产品推荐
相关产品推荐

