You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Unity MLAgents-learn无训练进展,Agent重复动作问题求助

问题分析与解决方案

核心问题汇总

  • Agent训练无进展,重复执行撞墙等相同动作
  • 撞墙触发负奖励后,GetCumulativeReward()显示总奖励未减少
  • 运行mlagents-learn约1分钟后崩溃,提示Aborted (core dumped)

问题排查与修复

1. 负奖励未生效的原因与修复

你的OnTriggerEnter中,负奖励分支调用EndEpisode()会直接重置Agent的奖励统计,而Debug日志在所有分支外,导致你看不到负奖励的临时累计效果。同时要确保碰撞体的Is Trigger已勾选,否则OnTriggerEnter不会触发。

修改代码,将日志放到分支内部:

private void OnTriggerEnter(Collider other){
    if (other.TryGetComponent<Goal>(out Goal goal)){
        AddReward(+1f);
        Debug.Log("到达终点,当前奖励: " + GetCumulativeReward());
        EndEpisode();
    }

    if (other.TryGetComponent<Wall>(out Wall wall)){
        AddReward(-1f);
        Debug.Log("撞墙,当前奖励: " + GetCumulativeReward());
        EndEpisode();
    }

    if(other.TryGetComponent<Car>(out Car car)){
        AddReward(-1f);
        Debug.Log("碰撞其他车辆,当前奖励: " + GetCumulativeReward());
        EndEpisode();
    }

    if(other.TryGetComponent<Checkpoint>(out Checkpoint checkpoint)){
        AddReward(+1f);
        Debug.Log("通过检查点,当前奖励: " + GetCumulativeReward());
    }
}

2. Agent训练无进展的关键修复

(1)补充观测信息

仅传入绝对位置不足以让模型判断动作合理性,需添加相对位置和自身速度:

private Rigidbody rb;

void Start(){
    rb = GetComponent<Rigidbody>();
    rb.useGravity = false; // 禁用重力,适配直道场景
}

public override void CollectObservations(VectorSensor sensor){
    // 目标与Agent的相对位置(比绝对位置更具指导意义)
    Vector3 relativePos = targetTransform.position - transform.position;
    sensor.AddObservation(relativePos);
    // Agent自身速度,让模型感知动作的物理反馈
    sensor.AddObservation(rb.velocity);
}

(2)改用物理驱动移动

直接修改transform.position跳过物理系统,模型无法学习动作的因果关系,改用Rigidbody控制:

public override void OnActionReceived(ActionBuffers actions){
    float moveX = actions.ContinuousActions[0];
    float moveZ = actions.ContinuousActions[1];

    float moveSpeed = 10f;
    // 用AddForce施加力,符合物理规律
    rb.AddForce(new Vector3(moveX, 0, moveZ) * moveSpeed);
    // 限制最大速度,避免Agent失控
    rb.velocity = Vector3.ClampMagnitude(rb.velocity, 15f);
}

(3)优化奖励函数

仅靠离散奖励无法引导Agent向目标移动,添加每帧的距离奖励:

private float lastDistance;

public override void OnEpisodeBegin(){
    transform.position = new Vector3(5.27417f, 0, -80.1366f);
    rb.velocity = Vector3.zero; // 重置速度
    // 记录初始与目标的距离
    lastDistance = Vector3.Distance(transform.position, targetTransform.position);
}

void Update(){
    float currentDistance = Vector3.Distance(transform.position, targetTransform.position);
    // 靠近目标给微小正奖励,远离给微小负奖励
    if(currentDistance < lastDistance){
        AddReward(0.01f);
    } else {
        AddReward(-0.005f);
    }
    lastDistance = currentDistance;
}

3. 崩溃问题Aborted (core dumped)解决

  • 控制内存占用:先以单Agent训练,避免多Agent同时运行导致内存溢出;减少不必要的观测维度。
  • 版本兼容:确保Unity版本与MLAgents包版本匹配(如Unity 2022对应MLAgents 2.0+),版本不兼容会引发底层崩溃。
  • 物理引擎异常:直接修改transform.position可能导致物理引擎状态紊乱,改用Rigidbody控制移动可缓解此问题。

额外建议

  • 启用TensorBoard可视化:运行mlagents-learn config.yaml --run-id=car_run后,用tensorboard --logdir=results查看奖励曲线,确认训练是否真的无进展。
  • 检查config.yaml:调整超参数(如学习率、批次大小),默认配置可能不匹配你的场景。

内容的提问来源于stack exchange,提问作者smitkims

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 05:01:47