You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在OpenAI Gym实现多智能体DDPG时Pygame窗口无响应崩溃

Pygame窗口无响应+CPU占用过高问题排查与修复(Boid集群DDPG实现)

核心问题定位

你的代码存在多个直接导致窗口崩溃、CPU拉满的问题,和算法(DQN/DDPG)无关,主要集中在Pygame事件处理、循环逻辑、维度匹配和代码冗余上:

  • Pygame事件队列未被处理,系统判定窗口无响应
  • 训练循环无终止条件,陷入无限循环
  • 状态输入与Actor网络维度不匹配,引发无效计算
  • 重复定义网络与超参数,造成内存冗余
  • Replay Buffer未填充,训练逻辑完全未执行但循环持续运行

修复步骤

1. 强制处理Pygame事件

Pygame需要定期处理系统事件(比如窗口关闭、鼠标操作),否则会被系统标记为无响应。在训练循环中加入事件处理:

while step_count < max_steps and not done:
    # 必须添加的事件处理
    for event in pygame.event.get():
        if event.type == pygame.QUIT:
            env.close()
            exit()
    # 原有训练逻辑...

2. 给训练循环添加终止条件

移除无限while True,设置最大步数或利用done标志终止循环:

total_episodes = 1000
max_steps_per_episode = 500  # 单回合最大步数
for episode in range(total_episodes):
    state = env.reset()
    episode_reward = 0
    step_count = 0
    done = False

    while step_count < max_steps_per_episode and not done:
        step_count += 1
        # 原有训练逻辑...

3. 删除重复的网络与超参数定义

代码中重复定义了两次超参数、Actor/Critic网络及优化器,直接删除第二份重复代码即可,保留第一份定义。

4. 修正状态输入与Actor网络维度匹配

你的observation_space是(50,4)的矩阵,但Actor网络的输入层只定义了state_size=4,维度完全不匹配。可以通过Flatten状态解决:

# 调整超参数中的state_size计算
state_size = env.observation_space.shape[0] * env.observation_space.shape[1]  # 50*4=200

# 训练循环中Flatten状态
state_flat = torch.tensor(state.flatten(), dtype=torch.float32)
action = actor(state_flat)
action = action.detach().numpy()

5. 填充Replay Buffer并执行训练更新

原代码未将交互数据加入Replay Buffer,导致ddpg_update永远不会执行,需要补充:

while step_count < max_steps_per_episode and not done:
    # ...原有逻辑(事件处理、动作生成)
    next_state, reward, done, _ = env.step(action)
    
    # 将交互数据加入Replay Buffer
    replay_buffer.append((state, action, reward, next_state, done))
    # 限制Buffer最大容量,避免内存溢出
    if len(replay_buffer) > buffer_size:
        replay_buffer.pop(0)
    
    # 执行DDPG更新
    ddpg_update()
    
    # 更新状态与奖励
    state = next_state
    episode_reward += reward
    
    # 调用渲染
    env.render()

6. 优化Boid邻居计算性能

get_closest_neighbors中对所有Boid排序是O(n²)操作,改用堆结构找最近邻居可降低计算负载:

import heapq
def get_closest_neighbors(self, boids, num_neighbors):
    distances = []
    for boid in boids:
        if boid is self:
            continue
        dist = self.position.distance_to(boid.position)
        distances.append((dist, boid))
    # 用heapq快速获取最小的N个邻居,比全排序高效
    closest = heapq.nsmallest(num_neighbors, distances, key=lambda x: x[0])
    neighbors = [b for d, b in closest]
    return neighbors

额外性能优化

  • 启用GPU加速:将PyTorch模型和张量移到GPU,减少CPU负载
    device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
    actor = Actor(state_size, action_size, hidden_size).to(device)
    target_actor = Actor(state_size, action_size, hidden_size).to(device)
    # 其余模型同理,训练时将张量传入device
    state_flat = torch.tensor(state.flatten(), dtype=torch.float32).to(device)
    
  • 降低渲染频率:比如每2步渲染一次,减少可视化对CPU的占用

内容的提问来源于stack exchange,提问作者user15752745

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 08:37:11