树莓派4运行深度强化学习代码报错求助
树莓派4上运行PPO深度强化学习任务报错的解决方法
问题背景
在树莓派4上部署控制步进电机的深度强化学习任务(基于Stable-Baselines3的PPO算法),代码在Colab环境可正常运行,但在树莓派上触发矩阵乘法相关的运行时错误。
报错信息
/home/pi/.local/lib/python3.9/site-packages/flatbuffers/compat.py:19: DeprecationWarning: the imp module is deprecated in favour of importlib; see the module's documentation for alternative uses import imp /home/pi/.local/lib/python3.9/site-packages/gym/spaces/box.py:128: UserWarning: WARN: Box bound precision lowered by casting to float32 logger.warn(f"Box bound precision lowered by casting to {self.dtype}") /home/pi/.local/lib/python3.9/site-packages/stable_baselines3/common/vec_env/patch_gym.py:49: UserWarning: You provided an OpenAI Gym environment. We strongly recommend transitioning to Gymnasium environments. Stable-Baselines3 is automatically wrapping your environments in a compatibility layer, which could potentially cause issues. warnings.warn( Using cpu device Traceback (most recent call last): File "/home/pi/reinf_nomotor.py", line 75, in <module> action, _ = model.predict(obs) File "/home/pi/.local/lib/python3.9/site-packages/stable_baselines3/common/base_class.py", line 555, in predict return self.policy.predict(observation, state, episode_start, deterministic) File "/home/pi/.local/lib/python3.9/site-packages/stable_baselines3/common/policies.py", line 349, in predict actions = self._predict(observation, deterministic=deterministic) File "/home/pi/.local/lib/python3.9/site-packages/stable_baselines3/common/policies.py", line 679, in _predict return self.get_distribution(observation).get_actions(deterministic=deterministic) File "/home/pi/.local/lib/python3.9/site-packages/stable_baselines3/common/policies.py", line 714, in get_distribution return self._get_action_dist_from_latent(latent_pi) File "/home/pi/.local/lib/python3.9/site-packages/stable_baselines3/common/policies.py", line 653, in _get_action_dist_from_latent mean_actions = self.action_net(latent_pi) File "/home/pi/.local/lib/python3.9/site-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl return self._call_impl(*args, **kwargs) File "/home/pi/.local/lib/python3.9/site-packages/torch/nn/modules/module.py", line 1527, in _call_impl return forward_call(*args, **kwargs) File "/home/pi/.local/lib/python3.9/site-packages/torch/nn/modules/linear.py", line 114, in forward return F.linear(input, self.weight, self.bias) RuntimeError: could not create a primitive descriptor for a matmul primitive
复现测试代码
import gym from stable_baselines3 import PPO from stable_baselines3.common.envs import DummyVecEnv import matplotlib.pyplot as plt import numpy as np # Define constants TARGET_DROP_SIZE = 25.0 ERROR_TOLERANCE = 0.01 # 1% error tolerance MAX_STEPS = 10 # Maximum number of steps per episode # Create a custom gym environment for drop size control class DropSizeControlEnv(gym.Env): def __init__(self): super(DropSizeControlEnv, self).__init__() # Environment parameters self.dt = 0.1 # Time step self.motor_angle = 90.0 # Initial motor angle # Action space: [-180, 180] self.action_space = gym.spaces.Box(low=np.array([-180.0]), high=np.array([180.0])) # Observation space: [0, 180] self.observation_space = gym.spaces.Box(low=np.array([0.0]), high=np.array([180.0])) def step(self, action): # Execute the control action (adjust motor angle) self.motor_angle += action[0] # Ensure the motor angle stays within bounds self.motor_angle = max(1.0, min(180.0, self.motor_angle)) # Simulate the achieved drop size (simplified for demonstration) achieved_drop_size = 24.5 + np.random.uniform(-0.5, 0.5) # Calculate the error error = TARGET_DROP_SIZE - achieved_drop_size # Calculate the reward (negative absolute error) reward = -abs(error) # Check if the error is within tolerance done = abs(error) <= ERROR_TOLERANCE # Return observation, reward, done, info return np.array([self.motor_angle]), reward, done, {} def reset(self): # Reset the environment to the initial state self.motor_angle = 90.0 return np.array([self.motor_angle]) # Create and wrap the custom environment env = DummyVecEnv([lambda: DropSizeControlEnv()]) # Create and train the PPO agent model = PPO("MlpPolicy", env, verbose=1) # Variables for tracking results time_steps = [] achieved_drop_sizes = [] target_drop_sizes = [] errors = [] # Training loop for episode in range(MAX_STEPS): obs = env.reset() while True: action, _ = model.predict(obs) obs, _, done, _ = env.step(action) # Simulate the achieved drop size achieved_drop_size = 24.5 + np.random.uniform(-0.5, 0.5) # Calculate the error error = TARGET_DROP_SIZE - achieved_drop_size # Store data for plotting time_steps.append(len(time_steps) * env.envs[0].dt) achieved_drop_sizes.append(achieved_drop_size) target_drop_sizes.append(TARGET_DROP_SIZE) errors.append(error) if done: break # Calculate the final achieved error final_error = abs(target_drop_sizes[-1] - achieved_drop_sizes[-1]) # Check if the achieved error is smaller than 1% if final_error < 0.01 * TARGET_DROP_SIZE: break # Plot the results plt.figure(figsize=(12, 6)) # Plot Achieved and Target Drop Size plt.subplot(1, 2, 1) plt.plot(time_steps, achieved_drop_sizes, label="Achieved Drop Size") plt.plot(time_steps, target_drop_sizes, label="Target Drop Size") plt.xlabel("Time (s)") plt.ylabel("Drop Size") plt.legend() plt.title("Drop Size Control") plt.grid(True) # Plot Error plt.subplot(1, 2, 2) plt.plot(time_steps, errors, label="Error") plt.xlabel("Time (s)") plt.ylabel("Error") plt.legend() plt.title("Error Plot") plt.grid(True) plt.tight_layout() plt.show()
解决方法
1. 降级PyTorch版本
树莓派的ARM架构对新版PyTorch的优化算子支持不完善,降级到兼容的稳定版本:
pip uninstall torch -y pip install torch==1.13.1+cpu torchvision==0.14.1+cpu --extra-index-url https://download.pytorch.org/whl/cpu
2. 禁用PyTorch的高级CPU优化
在代码开头添加以下配置,强制使用基础CPU算子,避免触发不兼容的矩阵乘法优化:
import torch # 限制线程数适配树莓派性能 torch.set_num_threads(1) # 禁用MKL-DNN和OpenMP优化 torch.backends.mkldnn.enabled = False torch.backends.openmp.enabled = False
3. 替换为Gymnasium环境
旧版Gym与Stable-Baselines3的兼容性层存在潜在问题,替换为官方推荐的Gymnasium:
pip uninstall gym -y pip install gymnasium
修改代码中的环境相关部分:
- 替换
import gym为import gymnasium as gym - 调整
step方法的返回格式,符合Gymnasium要求:return np.array([self.motor_angle]), reward, done, False, {} - 使用Stable-Baselines3针对Gymnasium的包装器:
from stable_baselines3.common.env_util import make_vec_env env = make_vec_env(lambda: DropSizeControlEnv(), n_envs=1)
4. 调整模型参数适配树莓派性能
降低模型复杂度,减少计算负载:
model = PPO( "MlpPolicy", env, verbose=1, policy_kwargs={"net_arch": [64, 32]}, # 缩小神经网络层数和神经元数量 device="cpu", # 明确指定使用CPU n_steps=64 # 减少每批次样本数量 )
内容的提问来源于stack exchange,提问作者Sk D
相关产品推荐
相关产品推荐

