在OpenAI Gym实现多智能体DDPG时Pygame窗口无响应崩溃
Pygame窗口无响应+CPU占用过高问题排查与修复(Boid集群DDPG实现)
核心问题定位
你的代码存在多个直接导致窗口崩溃、CPU拉满的问题,和算法(DQN/DDPG)无关,主要集中在Pygame事件处理、循环逻辑、维度匹配和代码冗余上:
- Pygame事件队列未被处理,系统判定窗口无响应
- 训练循环无终止条件,陷入无限循环
- 状态输入与Actor网络维度不匹配,引发无效计算
- 重复定义网络与超参数,造成内存冗余
- Replay Buffer未填充,训练逻辑完全未执行但循环持续运行
修复步骤
1. 强制处理Pygame事件
Pygame需要定期处理系统事件(比如窗口关闭、鼠标操作),否则会被系统标记为无响应。在训练循环中加入事件处理:
while step_count < max_steps and not done: # 必须添加的事件处理 for event in pygame.event.get(): if event.type == pygame.QUIT: env.close() exit() # 原有训练逻辑...
2. 给训练循环添加终止条件
移除无限while True,设置最大步数或利用done标志终止循环:
total_episodes = 1000 max_steps_per_episode = 500 # 单回合最大步数 for episode in range(total_episodes): state = env.reset() episode_reward = 0 step_count = 0 done = False while step_count < max_steps_per_episode and not done: step_count += 1 # 原有训练逻辑...
3. 删除重复的网络与超参数定义
代码中重复定义了两次超参数、Actor/Critic网络及优化器,直接删除第二份重复代码即可,保留第一份定义。
4. 修正状态输入与Actor网络维度匹配
你的observation_space是(50,4)的矩阵,但Actor网络的输入层只定义了state_size=4,维度完全不匹配。可以通过Flatten状态解决:
# 调整超参数中的state_size计算 state_size = env.observation_space.shape[0] * env.observation_space.shape[1] # 50*4=200 # 训练循环中Flatten状态 state_flat = torch.tensor(state.flatten(), dtype=torch.float32) action = actor(state_flat) action = action.detach().numpy()
5. 填充Replay Buffer并执行训练更新
原代码未将交互数据加入Replay Buffer,导致ddpg_update永远不会执行,需要补充:
while step_count < max_steps_per_episode and not done: # ...原有逻辑(事件处理、动作生成) next_state, reward, done, _ = env.step(action) # 将交互数据加入Replay Buffer replay_buffer.append((state, action, reward, next_state, done)) # 限制Buffer最大容量,避免内存溢出 if len(replay_buffer) > buffer_size: replay_buffer.pop(0) # 执行DDPG更新 ddpg_update() # 更新状态与奖励 state = next_state episode_reward += reward # 调用渲染 env.render()
6. 优化Boid邻居计算性能
get_closest_neighbors中对所有Boid排序是O(n²)操作,改用堆结构找最近邻居可降低计算负载:
import heapq def get_closest_neighbors(self, boids, num_neighbors): distances = [] for boid in boids: if boid is self: continue dist = self.position.distance_to(boid.position) distances.append((dist, boid)) # 用heapq快速获取最小的N个邻居,比全排序高效 closest = heapq.nsmallest(num_neighbors, distances, key=lambda x: x[0]) neighbors = [b for d, b in closest] return neighbors
额外性能优化
- 启用GPU加速:将PyTorch模型和张量移到GPU,减少CPU负载
device = torch.device("cuda" if torch.cuda.is_available() else "cpu") actor = Actor(state_size, action_size, hidden_size).to(device) target_actor = Actor(state_size, action_size, hidden_size).to(device) # 其余模型同理,训练时将张量传入device state_flat = torch.tensor(state.flatten(), dtype=torch.float32).to(device) - 降低渲染频率:比如每2步渲染一次,减少可视化对CPU的占用
内容的提问来源于stack exchange,提问作者user15752745
相关产品推荐
相关产品推荐

