Unity MLAgents-learn无训练进展,Agent重复动作问题求助
问题分析与解决方案
核心问题汇总
- Agent训练无进展,重复执行撞墙等相同动作
- 撞墙触发负奖励后,
GetCumulativeReward()显示总奖励未减少 - 运行
mlagents-learn约1分钟后崩溃,提示Aborted (core dumped)
问题排查与修复
1. 负奖励未生效的原因与修复
你的OnTriggerEnter中,负奖励分支调用EndEpisode()会直接重置Agent的奖励统计,而Debug日志在所有分支外,导致你看不到负奖励的临时累计效果。同时要确保碰撞体的Is Trigger已勾选,否则OnTriggerEnter不会触发。
修改代码,将日志放到分支内部:
private void OnTriggerEnter(Collider other){ if (other.TryGetComponent<Goal>(out Goal goal)){ AddReward(+1f); Debug.Log("到达终点,当前奖励: " + GetCumulativeReward()); EndEpisode(); } if (other.TryGetComponent<Wall>(out Wall wall)){ AddReward(-1f); Debug.Log("撞墙,当前奖励: " + GetCumulativeReward()); EndEpisode(); } if(other.TryGetComponent<Car>(out Car car)){ AddReward(-1f); Debug.Log("碰撞其他车辆,当前奖励: " + GetCumulativeReward()); EndEpisode(); } if(other.TryGetComponent<Checkpoint>(out Checkpoint checkpoint)){ AddReward(+1f); Debug.Log("通过检查点,当前奖励: " + GetCumulativeReward()); } }
2. Agent训练无进展的关键修复
(1)补充观测信息
仅传入绝对位置不足以让模型判断动作合理性,需添加相对位置和自身速度:
private Rigidbody rb; void Start(){ rb = GetComponent<Rigidbody>(); rb.useGravity = false; // 禁用重力,适配直道场景 } public override void CollectObservations(VectorSensor sensor){ // 目标与Agent的相对位置(比绝对位置更具指导意义) Vector3 relativePos = targetTransform.position - transform.position; sensor.AddObservation(relativePos); // Agent自身速度,让模型感知动作的物理反馈 sensor.AddObservation(rb.velocity); }
(2)改用物理驱动移动
直接修改transform.position跳过物理系统,模型无法学习动作的因果关系,改用Rigidbody控制:
public override void OnActionReceived(ActionBuffers actions){ float moveX = actions.ContinuousActions[0]; float moveZ = actions.ContinuousActions[1]; float moveSpeed = 10f; // 用AddForce施加力,符合物理规律 rb.AddForce(new Vector3(moveX, 0, moveZ) * moveSpeed); // 限制最大速度,避免Agent失控 rb.velocity = Vector3.ClampMagnitude(rb.velocity, 15f); }
(3)优化奖励函数
仅靠离散奖励无法引导Agent向目标移动,添加每帧的距离奖励:
private float lastDistance; public override void OnEpisodeBegin(){ transform.position = new Vector3(5.27417f, 0, -80.1366f); rb.velocity = Vector3.zero; // 重置速度 // 记录初始与目标的距离 lastDistance = Vector3.Distance(transform.position, targetTransform.position); } void Update(){ float currentDistance = Vector3.Distance(transform.position, targetTransform.position); // 靠近目标给微小正奖励,远离给微小负奖励 if(currentDistance < lastDistance){ AddReward(0.01f); } else { AddReward(-0.005f); } lastDistance = currentDistance; }
3. 崩溃问题Aborted (core dumped)解决
- 控制内存占用:先以单Agent训练,避免多Agent同时运行导致内存溢出;减少不必要的观测维度。
- 版本兼容:确保Unity版本与MLAgents包版本匹配(如Unity 2022对应MLAgents 2.0+),版本不兼容会引发底层崩溃。
- 物理引擎异常:直接修改
transform.position可能导致物理引擎状态紊乱,改用Rigidbody控制移动可缓解此问题。
额外建议
- 启用TensorBoard可视化:运行
mlagents-learn config.yaml --run-id=car_run后,用tensorboard --logdir=results查看奖励曲线,确认训练是否真的无进展。 - 检查
config.yaml:调整超参数(如学习率、批次大小),默认配置可能不匹配你的场景。
内容的提问来源于stack exchange,提问作者smitkims
相关产品推荐
相关产品推荐

