如何修复树莓派4摄像头流输入ResNet50时的图像形状错误?
问题定位与修复:ResNet50输入形状不匹配错误
错误代码行定位
错误出现在第53行:
predictions = classifier.predict(np.expand_dims(image, axis=-1))
问题根源分析
- 格式不匹配:你配置的lores流是
YUV420格式,capture_buffer("lores")获取的是单通道灰度数据(形状为(240, 320)),直接转成PIL Image后仍是单通道,而非ResNet50要求的3通道RGB图像。 - 多余维度扩展:你先通过
np.expand_dims(image, axis=0)添加了batch维度(形状变为(1, 224, 224)),又额外执行了一次np.expand_dims(image, axis=-1),最终得到(1, 224, 224, 1)的单通道输入,与ResNet50要求的(None, 224, 224, 3)形状完全不符。
修复方案
核心修改点
- 将YUV420格式转换为RGB三通道图像
- 移除多余的维度扩展操作
- 使用ResNet官方预处理函数替代手动归一化
修改后的循环代码段
while True: # Capture a frame from the low-resolution stream cur = picam2.capture_buffer("lores") # 处理YUV420转RGB:拆分Y/U/V通道并转换格式 y = cur[:w*h].reshape(h, w) u = cur[w*h:w*h + (w//2)*(h//2)].reshape(h//2, w//2) v = cur[w*h + (w//2)*(h//2):].reshape(h//2, w//2) # 将U/V通道放大至与Y通道同尺寸 u = np.repeat(np.repeat(u, 2, axis=0), 2, axis=1) v = np.repeat(np.repeat(v, 2, axis=0), 2, axis=1) # 按BT.601标准完成YUV到RGB的转换 yuv = np.stack([y, u, v], axis=-1).astype(np.float32) yuv[:, :, 0] -= 16 yuv[:, :, 1:] -= 128 conversion_matrix = np.array([[1.164, 0.000, 1.596], [1.164, -0.392, -0.813], [1.164, 2.017, 0.000]]) rgb = np.dot(yuv, conversion_matrix.T) rgb = np.clip(rgb, 0, 255).astype(np.uint8) # 预处理图像以匹配ResNet50要求 image = Image.fromarray(rgb) image = image.resize((224, 224)) # 使用官方预处理函数,替代手动归一化 image = tf.keras.applications.resnet50.preprocess_input(np.array(image)) image = np.expand_dims(image, axis=0) # 仅添加batch维度 # 执行模型预测 predictions = classifier.predict(image) # 打印预测结果(取Top1分类) print("Model predictions class:", tf.keras.applications.imagenet_utils.decode_predictions(preds=predictions, top=1)[0][0][1]) time.sleep(1)
补充说明
- ResNet50的输入预处理并非简单除以255,官方
preprocess_input函数会基于ImageNet数据集的均值和标准差做标准化,能提升分类准确性。 - 必须完成YUV到RGB的转换,否则即使强行调整形状为3通道,单通道灰度图也无法被模型正确识别。
内容的提问来源于stack exchange,提问作者Skywalker
相关产品推荐
相关产品推荐

