You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Mujoco sim.step()执行后观测值未更新问题求助

强化学习循环中观测值未更新问题排查

问题描述

我有一段非循环Python代码可正常运行:

sim.reset()
sim.step()
print(sim.data.qpos)
print(sim.data.get_body_xpos("EE"))

print()
sim.data.ctrl[:] = [-0.244,-0.66,-0.785,-1.79,2.42,-0.126]
sim.step()
print(sim.get_state())
print(sim.data.get_body_xpos("EE"))

print()
sim.data.qpos[:] = [-0.244,-0.66,-0.785,-1.79,2.42,-0.126]
sim.forward()
sim.step()
print(sim.data.qpos)
print(sim.data.get_body_xpos("EE"))

但在强化学习循环代码中,同一时间步内的observation与new_observation值完全相同(尽管每次action不同):

EPOCHS = 5000
STEPS = 100000
TARGET_POS = [-0.47367866, 0.01074746, 0.76153706]


for i in range(EPOCHS):
    done = 0
    sim.reset()
    sim.forward()
    sim.step()
    observation = sim.data.get_body_xpos("EE")
    distance = 0
    iterator = 0
    while not done:
        action = agent.select_action(observation)
        sim.data.ctrl[:] = action
        sim.forward()
        sim.step() #required to update the joints and EE

        new_observation = sim.data.get_body_xpos("EE")
      
        print(iterator, observation)
        print(iterator, new_observation)
     
        reward = ur5_reward_func(new_observation, TARGET_POS)
        distance += reward

        iterator+=1

        done = check_in_target_sphere(TARGET_POS,new_observation) #function returns a 1 if the EE coordinates are in the sphere, else 0
        
        agent.replay_buffer.append(observation, action, reward, done, new_observation, iterator)

        if( agent.replay_buffer.get_buffer_size() > agent.REPLAY_BATCH_SIZE ):
            agent.update_parameters()
        
        

        observation = new_observation

        if(done == True):
            break

部分输出示例:

0 [-1.91900000e-01 -3.16096174e-08  1.00110000e+00]
0 [-1.91900000e-01 -3.16096174e-08  1.00110000e+00]
1 [-1.91900000e-01 -1.07517997e-06  1.00110000e+00]
1 [-1.91900000e-01 -1.07517997e-06  1.00110000e+00]

错误原因分析

问题出在多余的sim.forward()调用:

  • 在MuJoCo中,sim.step()方法内部已经包含完整的动力学更新流程:应用控制信号 → 执行动力学积分 → 计算正向运动学(更新关节、末端执行器位置等)。
  • 你在设置sim.data.ctrl[:] = action后先调用sim.forward(),此时正向运动学计算基于的是上一步的关节状态,并未应用当前的控制信号;后续的sim.step()虽然会更新状态,但重复调用forward干扰了正常的状态更新逻辑,导致观测值未按预期变化。
  • 另外,sim.reset()后的sim.forward()也是多余的,sim.reset()已经将状态重置到初始值,直接调用sim.step()或仅sim.forward()即可获取初始观测。

修正后的代码

核心修改点:

  1. 移除设置控制信号后的sim.forward()调用
  2. 优化初始状态获取流程
EPOCHS = 5000
STEPS = 100000
TARGET_POS = [-0.47367866, 0.01074746, 0.76153706]


for i in range(EPOCHS):
    done = 0
    sim.reset()
    # 重置后直接step获取初始状态,或仅调用sim.forward()获取初始观测(根据需求选择)
    sim.step()
    observation = sim.data.get_body_xpos("EE")
    distance = 0
    iterator = 0
    while not done:
        action = agent.select_action(observation)
        sim.data.ctrl[:] = action
        # 移除多余的sim.forward(),直接调用step即可完成状态更新
        sim.step()

        new_observation = sim.data.get_body_xpos("EE")
      
        print(iterator, observation)
        print(iterator, new_observation)
     
        reward = ur5_reward_func(new_observation, TARGET_POS)
        distance += reward

        iterator+=1

        done = check_in_target_sphere(TARGET_POS,new_observation)
        
        agent.replay_buffer.append(observation, action, reward, done, new_observation, iterator)

        if agent.replay_buffer.get_buffer_size() > agent.REPLAY_BATCH_SIZE:
            agent.update_parameters()
        
        observation = new_observation

        if done:
            break

额外建议

如果发现sim.data.get_body_xpos返回的是数组引用导致观测值被意外覆盖,可以将观测值转为副本存储:

# 获取观测时转为副本
observation = sim.data.get_body_xpos("EE").copy()
new_observation = sim.data.get_body_xpos("EE").copy()

内容的提问来源于stack exchange,提问作者Patrick Adjei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 05:45:10