Windows正常运行的ML Agents项目在Linux执行启动失败求助
ML Agents Linux远程启动失败问题排查方案
环境信息
- Python 3.8.10
- ml-agents: 0.27.0
- ml-agents-envs: 0.27.0
- Communicator API: 1.5.0
- PyTorch: 1.8.1+cu102
- Unity版本:2022.3.2f1(曾尝试2020.2.6f1,无改善)
- 目标系统:Ubuntu 20.04.6 LTS(SSH远程运行)
问题场景
Windows本地开发运行ML Agents仿真项目完全正常,切换到Ubuntu服务器提升仿真效率,将项目打包为Linux可执行文件后,通过SSH远程执行训练命令或直接用Python调用UnityEnvironment均失败,报错返回码1。
操作流程
- Unity打包步骤:
File->Build Settings->Windows, Mac, Linux->Linux,勾选Development Build后执行Build - 将打包后的文件上传至Ubuntu服务器
- 执行训练命令:
mlagents-learn config/ppoagent.yaml --env=visibility_game_linux_build2022.x86_64 --no-graphics --force
- 或尝试Python代码直接启动环境:
from mlagents_envs.environment import UnityEnvironment env = UnityEnvironment(file_name='visibility_game_linux_build2022.x86_64')
报错信息
训练命令报错
[INFO] Learning was interrupted. Please wait while the graph is generated. Traceback (most recent call last): File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/bin/mlagents-learn", line 8, in sys.exit(main()) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents/trainers/learn.py", line 250, in main run_cli(parse_command_line()) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents/trainers/learn.py", line 246, in run_cli run_training(run_seed, options) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents/trainers/learn.py", line 125, in run_training tc.start_learning(env_manager) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents_envs/timers.py", line 305, in wrapped return func(*args, **kwargs) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents/trainers/trainer_controller.py", line 198, in start_learning raise ex File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents/trainers/trainer_controller.py", line 173, in start_learning self._reset_env(env_manager) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents_envs/timers.py", line 305, in wrapped return func(*args, **kwargs) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents/trainers/trainer_controller.py", line 105, in _reset_env env_manager.reset(config=new_config) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents/trainers/env_manager.py", line 68, in reset self.first_step_infos = self._reset_env(config) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents/trainers/subprocess_env_manager.py", line 334, in _reset_env ew.previous_step = EnvironmentStep(ew.recv().payload, ew.worker_id, {}, {}) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents/trainers/subprocess_env_manager.py", line 98, in recv raise env_exception mlagents_envs.exception.UnityEnvironmentException: Environment shut down with return code 1.
Python代码启动报错
Traceback (most recent call last): File "", line 1, in File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents_envs/environment.py", line 223, in init aca_output = self._send_academy_parameters(rl_init_parameters_in) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents_envs/environment.py", line 477, in _send_academy_parameters return self._communicator.initialize(inputs, self._poll_process) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents_envs/rpc_communicator.py", line 121, in initialize self.poll_for_timeout(poll_callback) File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents_envs/rpc_communicator.py", line 108, in poll_for_timeout poll_callback() File "/home/rmarr/Documents/GflowsForSimulation_env/GflowsForSimulation_venv_real/lib/python3.8/site-packages/mlagents_envs/environment.py", line 403, in _poll_process raise UnityEnvironmentException(exc_msg) mlagents_envs.exception.UnityEnvironmentException: Environment shut down with return code 1.
排查与解决步骤
添加可执行权限
确保Linux可执行文件有运行权限:chmod +x visibility_game_linux_build2022.x86_64检查依赖缺失
直接运行可执行文件,查看具体依赖报错:./visibility_game_linux_build2022.x86_64常见缺失依赖及安装命令:
sudo apt-get install libgconf-2-4 libgtk-3-0 libx11-xcb1 libxcb-dri3-0 libnss3 libasound2无图形界面服务器需安装虚拟显示驱动:
sudo apt-get install xvfb用xvfb启动训练:
xvfb-run -a mlagents-learn config/ppoagent.yaml --env=visibility_game_linux_build2022.x86_64 --no-graphics --force修正Unity打包配置
- 确认打包架构为Linux x86_64
- 取消勾选Unity Remote Support
- 移除项目中Windows特定代码(如
System.Windows相关API),替换为跨平台实现
版本兼容性校验
尝试升级ml-agents到适配Unity 2022.x的0.28.x版本,或降级到对应稳定版本,确保Communicator API版本完全匹配设置显示环境变量
在SSH会话中添加环境变量:export DISPLAY=:0.0或直接在启动命令前追加:
DISPLAY=:0.0 mlagents-learn config/ppoagent.yaml --env=visibility_game_linux_build2022.x86_64 --no-graphics --force查看Unity崩溃日志
日志路径通常为~/.config/unity3d/[你的公司名称]/[项目名称]/Player.log,通过日志定位具体崩溃原因(如资源缺失、脚本错误等)
内容的提问来源于stack exchange,提问作者Ryan Marr
相关产品推荐
相关产品推荐

