You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Abseil环境下OpenAI Gym运行Super Mario Bros多进程的问题求助

问题

使用Abseil库结合OpenAI Gym进行多进程开发时,普通gym.make创建的环境(如LunarLander-v2)可正常运行,但gym-super-mario-bros创建的Super Mario Bros环境无法工作。

最小复现代码

from absl import app
import os
os.environ['OMP_NUM_THREADS'] = '1'

import gym
import gym_super_mario_bros
from gym_super_mario_bros.actions import SIMPLE_MOVEMENT
from nes_py.wrappers import JoypadSpace
import multiprocessing as mp
import torch
from torch import nn
import time


def get_env():
    env = JoypadSpace(gym_super_mario_bros.make('SuperMarioBros-1-1-v0'), SIMPLE_MOVEMENT)
    # env = gym.make('LunarLander-v2') # other environment such as this one and others works well
    return env

def do_something(env, net1, net2):
    print('inside do_something')
    obs = env.reset()
    print(f'after reset {obs.shape}')
    net2.load_state_dict(net1.state_dict())
    print('after load_state_dict')

def main(args):
    del args
    env = get_env()

    net1 = nn.Sequential(nn.Conv2d(1, 20, 5), nn.ReLU())
    net2 = nn.Sequential(nn.Conv2d(1, 20, 5), nn.ReLU())

    net1.share_memory()
    net2.share_memory()

    device = torch.device('cuda')
    net1 = net1.to(device)
    net2 = net2.to(device)

    p = mp.Process(target=do_something, args=(env, net1, net2,))
    p.start()

    time.sleep(4.0) # wait for the above process to execute print statements
    env.close()

if __name__ == '__main__':
    mp.set_start_method('spawn')
    app.run(main)

运行报错信息

$ python mwe.py 
inside do_something
terminate called after throwing an instance of 'std::bad_alloc'
  what():  std::bad_alloc
[W CudaIPCTypes.cpp:15] Producer process has been terminated before all shared CUDA tensors released. See Note [Sharing CUDA tensors]
[W CUDAGuardImpl.h:46] Warning: CUDA warning: driver shutting down (function uncheckedGetDevice)
[W CUDAGuardImpl.h:62] Warning: CUDA warning: invalid device ordinal (function uncheckedSetDevice)
[W CUDAGuardImpl.h:46] Warning: CUDA warning: driver shutting down (function uncheckedGetDevice)
[W CUDAGuardImpl.h:62] Warning: CUDA warning: invalid device ordinal (function uncheckedSetDevice)
[W CUDAGuardImpl.h:46] Warning: CUDA warning: driver shutting down (function uncheckedGetDevice)
[W CUDAGuardImpl.h:62] Warning: CUDA warning: invalid device ordinal (function uncheckedSetDevice)
[W CUDAGuardImpl.h:46] Warning: CUDA warning: driver shutting down (function uncheckedGetDevice)
[W CUDAGuardImpl.h:62] Warning: CUDA warning: invalid device ordinal (function uncheckedSetDevice)

环境版本信息

LibraryVersion
absl-py1.3.0
cuda11.7
gym0.17.2
gym-super-mario-bros7.4.0
nes-py8.2.1
numpy1.21.0
python3.9.16
torch1.13.1

解决方案

1. 子进程内单独初始化环境

gym-super-mario-bros依赖的nes-py底层使用C扩展,无法跨进程安全共享环境实例。修改代码,让每个子进程自行创建环境:

from absl import app
import os
os.environ['OMP_NUM_THREADS'] = '1'

import gym
import gym_super_mario_bros
from gym_super_mario_bros.actions import SIMPLE_MOVEMENT
from nes_py.wrappers import JoypadSpace
import multiprocessing as mp
import torch
from torch import nn
import time


def get_env():
    env = JoypadSpace(gym_super_mario_bros.make('SuperMarioBros-1-1-v0'), SIMPLE_MOVEMENT)
    return env

def do_something(net1_state_dict, device_id):
    print('inside do_something')
    # 子进程内创建环境
    env = get_env()
    obs = env.reset()
    print(f'after reset {obs.shape}')
    
    # 子进程内初始化网络并加载参数
    net2 = nn.Sequential(nn.Conv2d(1, 20, 5), nn.ReLU())
    net2.to(torch.device(f'cuda:{device_id}'))
    net2.load_state_dict(net1_state_dict)
    print('after load_state_dict')
    
    env.close()

def main(args):
    del args

    net1 = nn.Sequential(nn.Conv2d(1, 20, 5), nn.ReLU())
    device = torch.device('cuda')
    net1 = net1.to(device)

    # 传递网络状态字典而非共享CUDA张量
    net1_state_dict = net1.state_dict()

    p = mp.Process(target=do_something, args=(net1_state_dict, device.index,))
    p.start()
    p.join() # 用join替代sleep,确保子进程执行完毕

if __name__ == '__main__':
    mp.set_start_method('spawn')
    app.run(main)

2. 避免跨进程共享CUDA张量

原代码直接传递CUDA张量给子进程会引发CUDA IPC问题,改为传递网络状态字典,在子进程内重新初始化网络并加载参数,同时指定正确的CUDA设备。

3. 确保子进程独立管理资源

使用spawn启动方式时,子进程会重新导入模块并初始化,所有依赖C扩展的资源(如NES模拟器环境)必须在子进程内单独创建,不能从主进程传递。

内容的提问来源于stack exchange,提问作者ravi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 20:17:53