You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

树莓派4运行深度强化学习代码报错求助

树莓派4上运行PPO深度强化学习任务报错的解决方法

问题背景

在树莓派4上部署控制步进电机的深度强化学习任务(基于Stable-Baselines3的PPO算法),代码在Colab环境可正常运行,但在树莓派上触发矩阵乘法相关的运行时错误。

报错信息

/home/pi/.local/lib/python3.9/site-packages/flatbuffers/compat.py:19: DeprecationWarning: the imp module is deprecated in favour of importlib; see the module's documentation for alternative uses
  import imp
/home/pi/.local/lib/python3.9/site-packages/gym/spaces/box.py:128: UserWarning: WARN: Box bound precision lowered by casting to float32
  logger.warn(f"Box bound precision lowered by casting to {self.dtype}")
/home/pi/.local/lib/python3.9/site-packages/stable_baselines3/common/vec_env/patch_gym.py:49: UserWarning: You provided an OpenAI Gym environment. We strongly recommend transitioning to Gymnasium environments. Stable-Baselines3 is automatically wrapping your environments in a compatibility layer, which could potentially cause issues.
  warnings.warn(
Using cpu device
Traceback (most recent call last):
  File "/home/pi/reinf_nomotor.py", line 75, in <module>
    action, _ = model.predict(obs)
  File "/home/pi/.local/lib/python3.9/site-packages/stable_baselines3/common/base_class.py", line 555, in predict
    return self.policy.predict(observation, state, episode_start, deterministic)
  File "/home/pi/.local/lib/python3.9/site-packages/stable_baselines3/common/policies.py", line 349, in predict
    actions = self._predict(observation, deterministic=deterministic)
  File "/home/pi/.local/lib/python3.9/site-packages/stable_baselines3/common/policies.py", line 679, in _predict
    return self.get_distribution(observation).get_actions(deterministic=deterministic)
  File "/home/pi/.local/lib/python3.9/site-packages/stable_baselines3/common/policies.py", line 714, in get_distribution
    return self._get_action_dist_from_latent(latent_pi)
  File "/home/pi/.local/lib/python3.9/site-packages/stable_baselines3/common/policies.py", line 653, in _get_action_dist_from_latent
    mean_actions = self.action_net(latent_pi)
  File "/home/pi/.local/lib/python3.9/site-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl
    return self._call_impl(*args, **kwargs)
  File "/home/pi/.local/lib/python3.9/site-packages/torch/nn/modules/module.py", line 1527, in _call_impl
    return forward_call(*args, **kwargs)
  File "/home/pi/.local/lib/python3.9/site-packages/torch/nn/modules/linear.py", line 114, in forward
    return F.linear(input, self.weight, self.bias)
RuntimeError: could not create a primitive descriptor for a matmul primitive

复现测试代码

import gym
from stable_baselines3 import PPO
from stable_baselines3.common.envs import DummyVecEnv
import matplotlib.pyplot as plt
import numpy as np

# Define constants
TARGET_DROP_SIZE = 25.0
ERROR_TOLERANCE = 0.01  # 1% error tolerance
MAX_STEPS = 10  # Maximum number of steps per episode

# Create a custom gym environment for drop size control
class DropSizeControlEnv(gym.Env):
    def __init__(self):
        super(DropSizeControlEnv, self).__init__()

        # Environment parameters
        self.dt = 0.1  # Time step
        self.motor_angle = 90.0  # Initial motor angle

        # Action space: [-180, 180]
        self.action_space = gym.spaces.Box(low=np.array([-180.0]), high=np.array([180.0]))
        # Observation space: [0, 180]
        self.observation_space = gym.spaces.Box(low=np.array([0.0]), high=np.array([180.0]))

    def step(self, action):
        # Execute the control action (adjust motor angle)
        self.motor_angle += action[0]

        # Ensure the motor angle stays within bounds
        self.motor_angle = max(1.0, min(180.0, self.motor_angle))

        # Simulate the achieved drop size (simplified for demonstration)
        achieved_drop_size = 24.5 + np.random.uniform(-0.5, 0.5)

        # Calculate the error
        error = TARGET_DROP_SIZE - achieved_drop_size

        # Calculate the reward (negative absolute error)
        reward = -abs(error)

        # Check if the error is within tolerance
        done = abs(error) <= ERROR_TOLERANCE

        # Return observation, reward, done, info
        return np.array([self.motor_angle]), reward, done, {}

    def reset(self):
        # Reset the environment to the initial state
        self.motor_angle = 90.0
        return np.array([self.motor_angle])

# Create and wrap the custom environment
env = DummyVecEnv([lambda: DropSizeControlEnv()])

# Create and train the PPO agent
model = PPO("MlpPolicy", env, verbose=1)

# Variables for tracking results
time_steps = []
achieved_drop_sizes = []
target_drop_sizes = []
errors = []

# Training loop
for episode in range(MAX_STEPS):
    obs = env.reset()
    while True:
        action, _ = model.predict(obs)
        obs, _, done, _ = env.step(action)

        # Simulate the achieved drop size
        achieved_drop_size = 24.5 + np.random.uniform(-0.5, 0.5)

        # Calculate the error
        error = TARGET_DROP_SIZE - achieved_drop_size

        # Store data for plotting
        time_steps.append(len(time_steps) * env.envs[0].dt)
        achieved_drop_sizes.append(achieved_drop_size)
        target_drop_sizes.append(TARGET_DROP_SIZE)
        errors.append(error)

        if done:
            break

    # Calculate the final achieved error
    final_error = abs(target_drop_sizes[-1] - achieved_drop_sizes[-1])

    # Check if the achieved error is smaller than 1%
    if final_error < 0.01 * TARGET_DROP_SIZE:
        break

# Plot the results
plt.figure(figsize=(12, 6))

# Plot Achieved and Target Drop Size
plt.subplot(1, 2, 1)
plt.plot(time_steps, achieved_drop_sizes, label="Achieved Drop Size")
plt.plot(time_steps, target_drop_sizes, label="Target Drop Size")
plt.xlabel("Time (s)")
plt.ylabel("Drop Size")
plt.legend()
plt.title("Drop Size Control")
plt.grid(True)

# Plot Error
plt.subplot(1, 2, 2)
plt.plot(time_steps, errors, label="Error")
plt.xlabel("Time (s)")
plt.ylabel("Error")
plt.legend()
plt.title("Error Plot")
plt.grid(True)

plt.tight_layout()
plt.show()

解决方法

1. 降级PyTorch版本

树莓派的ARM架构对新版PyTorch的优化算子支持不完善,降级到兼容的稳定版本:

pip uninstall torch -y
pip install torch==1.13.1+cpu torchvision==0.14.1+cpu --extra-index-url https://download.pytorch.org/whl/cpu

2. 禁用PyTorch的高级CPU优化

在代码开头添加以下配置,强制使用基础CPU算子,避免触发不兼容的矩阵乘法优化:

import torch
# 限制线程数适配树莓派性能
torch.set_num_threads(1)
# 禁用MKL-DNN和OpenMP优化
torch.backends.mkldnn.enabled = False
torch.backends.openmp.enabled = False

3. 替换为Gymnasium环境

旧版Gym与Stable-Baselines3的兼容性层存在潜在问题,替换为官方推荐的Gymnasium:

pip uninstall gym -y
pip install gymnasium

修改代码中的环境相关部分:

  • 替换import gym为import gymnasium as gym
  • 调整step方法的返回格式,符合Gymnasium要求:
    return np.array([self.motor_angle]), reward, done, False, {}
    
  • 使用Stable-Baselines3针对Gymnasium的包装器:
    from stable_baselines3.common.env_util import make_vec_env
    env = make_vec_env(lambda: DropSizeControlEnv(), n_envs=1)
    

4. 调整模型参数适配树莓派性能

降低模型复杂度,减少计算负载:

model = PPO(
    "MlpPolicy", 
    env, 
    verbose=1,
    policy_kwargs={"net_arch": [64, 32]},  # 缩小神经网络层数和神经元数量
    device="cpu",  # 明确指定使用CPU
    n_steps=64  # 减少每批次样本数量
)

内容的提问来源于stack exchange,提问作者Sk D

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 08:08:17