Unity ML-Agents夺旗游戏Agent瞬移/极速移动异常排查
Unity ML-Agents夺旗游戏Agent异常行为问题
当前游戏设置
- 2个Agent(每队1个)
- 每队各有1面旗帜(flag)和1个基地(base)
- 地面已添加NavMesh,Agent挂载NavMesh Agent组件
- 游戏目标:Agent夺取对方旗帜并返回己方基地
问题现象
运行模拟时,Agent出现以下异常行为:
- 瞬间瞬移到旗帜附近的特定位置(无过渡移动过程,速度极快)
- 随后停留在该位置直到回合结束
- 完全没有正常移动动画,仿佛直接跳转至目标点
怀疑问题出在velocity设置或对策略返回的ActionBuffers的解析上,以下是代码简化版本:
public override void CollectObservations(VectorSensor sensor) { // 3 + 3 = 6 floats sensor.AddObservation(transform.localPosition); sensor.AddObservation(flag.localPosition); sensor.AddObservation(myBase.localPosition); } public override void OnActionReceived(ActionBuffers actions) { float moveThreshold = 0.1f; // 添加移动阈值 Vector3 moveInput = new Vector3(actions.ContinuousActions[0], 0, actions.ContinuousActions[1]); // 只有当输入超过阈值时才移动 if (moveInput.magnitude > moveThreshold) { Vector3 move = moveInput.normalized * moveSpeed; rb.velocity = move; } else { rb.velocity = Vector3.zero; } AddReward(MaxStep > 0 ? -1f / MaxStep : -0.001f); } public override void Heuristic(in ActionBuffers a) { var c = a.ContinuousActions; c[0] = Input.GetAxis("Horizontal"); c[1] = Input.GetAxis("Vertical"); } void OnTriggerEnter(Collider other) { if (other.CompareTag("Flag")) { AddReward(1f); // 捡起对方旗帜 other.gameObject.SetActive(false); } if (other.CompareTag("Base") && !other.gameObject.CompareTag(gameObject.tag)) { AddReward(2f); // 将旗帜带回己方基地——获胜 EndEpisode(); } } public override void OnEpisodeBegin() { // 随机初始位置 transform.localPosition = new Vector3(Random.Range(-4, 4), 0.5f, Random.Range(-4, 4)); flag.gameObject.SetActive(true); rb.linearVelocity = Vector3.zero; rb.angularVelocity = Vector3.zero; }
相关截图
问题排查与解决方案
1. 核心冲突:NavMesh Agent与Rigidbody共存
你同时给Agent挂载了NavMesh Agent和Rigidbody组件,二者都会控制物体移动,直接导致逻辑冲突——NavMesh Agent可能强制Agent瞬移到导航目标点,随后被Rigidbody的velocity设置锁死在原地。
解决办法:
二选一禁用其中一个组件:
- 保留NavMesh Agent:删除Rigidbody,在
OnActionReceived中改为通过navMeshAgent.destination或navMeshAgent.velocity控制移动,不要操作Rigidbody。 - 保留Rigidbody:移除或禁用NavMesh Agent组件,完全通过物理逻辑控制移动。
2. 移动逻辑优化
当前的移动阈值0.1f过高,可能过滤掉策略输出的微小移动指令,导致Agent瞬移后无法继续移动;同时直接设置rb.velocity会覆盖物理系统的自然减速效果。
优化代码示例:
public override void OnActionReceived(ActionBuffers actions) { Vector3 moveInput = new Vector3(actions.ContinuousActions[0], 0, actions.ContinuousActions[1]); if (moveInput.magnitude > 0.01f) { Vector3 moveDir = moveInput.normalized; // 用AddForce替代直接设置velocity,保留物理特性 rb.AddForce(moveDir * moveSpeed, ForceMode.Force); // 限制最大移动速度 if (rb.velocity.magnitude > moveSpeed) { rb.velocity = rb.velocity.normalized * moveSpeed; } } else { // 平滑减速,避免突然停住 rb.velocity = Vector3.Lerp(rb.velocity, Vector3.zero, Time.deltaTime * 5f); } AddReward(MaxStep > 0 ? -1f / MaxStep : -0.001f); }
3. 观测与奖励逻辑补全
当前观测缺少关键信息,导致策略学习不完整:
- 补充是否携带旗帜的状态:
// 在CollectObservations中添加 sensor.AddObservation(isHoldingFlag ? 1f : 0f); - 若存在对方Agent,补充其位置信息,帮助策略做出更合理的决策。
同时确保获胜奖励足够突出,避免Agent陷入“停滞减少惩罚”的局部最优。
4. 场景重置校验
检查OnEpisodeBegin中是否完整重置所有状态:
- 旗帜不仅要设置为active,还要重置到初始位置
- 确认Agent的随机初始位置在NavMesh范围内,避免导航异常
内容的提问来源于stack exchange,提问作者Avi Garg
相关产品推荐
相关产品推荐




