如何用Keras+OpenCV实现摄像头实时图像识别并显示标签与准确率
解决方案:实时摄像头图像识别(Keras + OpenCV)
Got it, let's get your real-time camera recognition up and running step by step! I'll adapt your existing local image test code to work with live camera feeds, and explain all the key parts so you understand how it loops and displays predictions.
整体实现思路
We'll use OpenCV's VideoCapture to access your camera, then run an infinite loop that:
- Reads each frame from the camera
- Applies the exact same preprocessing you used for local images
- Feeds the processed frame into your Keras model for prediction
- Draws the predicted label directly onto the live video feed
- Lets you exit the loop with a key press
完整代码实现
import cv2 import numpy as np from keras import backend as K # 复用你提供的分类配置 num_classes = 4 labels = np.zeros(808) # 补全你原本的标签初始化逻辑 labels[0:202] = 0 labels[202:404] = 1 labels[404:606] = 2 labels[606:] = 3 names = ['cats', 'dogs', 'horses', 'humans'] num_channel = 1 # 从你的测试代码看是单通道灰度图,务必和训练时一致 # 加载你训练好的模型(记得替换成你的模型路径) # model = load_model('your_trained_model.h5') # 初始化摄像头:0代表默认摄像头(多摄像头可尝试1、2等) cap = cv2.VideoCapture(0) # 检查摄像头是否成功启动 if not cap.isOpened(): print("摄像头无法打开,请检查设备连接或权限!") exit() # 启动实时识别循环 while True: # 读取一帧图像:ret表示是否读取成功,frame是图像本身 ret, frame = cap.read() if not ret: print("无法获取图像帧,退出循环") break # 对帧执行和本地测试完全一致的预处理 processed_frame = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY) processed_frame = cv2.resize(processed_frame, (128, 128)) processed_frame = np.array(processed_frame) processed_frame = processed_frame.astype('float32') processed_frame /= 255 # 扩展维度以匹配模型输入格式 if num_channel == 1: if K.image_dim_ordering() == 'th': processed_frame = np.expand_dims(processed_frame, axis=0) processed_frame = np.expand_dims(processed_frame, axis=0) else: processed_frame = np.expand_dims(processed_frame, axis=3) processed_frame = np.expand_dims(processed_frame, axis=0) else: if K.image_dim_ordering() == 'th': processed_frame = np.rollaxis(processed_frame, 2, 0) processed_frame = np.expand_dims(processed_frame, axis=0) else: processed_frame = np.expand_dims(processed_frame, axis=0) # 执行预测并获取类别名称 pred_class = model.predict_classes(processed_frame)[0] pred_name = names[pred_class] # 在实时画面上绘制预测标签 cv2.putText(frame, f"Prediction: {pred_name}", (10, 30), cv2.FONT_HERSHEY_SIMPLEX, 1, (0, 255, 0), 2) # 显示带标签的实时画面 cv2.imshow('Real-Time Animal/Human Recognition', frame) # 按下小写q键退出循环(确保窗口处于激活状态) if cv2.waitKey(1) & 0xFF == ord('q'): break # 清理资源:释放摄像头并关闭所有窗口 cap.release() cv2.destroyAllWindows()
关键细节说明
- 摄像头初始化:
cv2.VideoCapture(0)targets your default webcam. If you have multiple cameras, try passing 1, 2, etc., to switch. - 循环逻辑:The
while Trueloop runs until you press 'q' or the camera disconnects.cap.read()grabs each frame in real time. - 预处理一致性:This is critical! The preprocessing steps (grayscale conversion, resizing, normalization) must match exactly what you used to train your model—otherwise predictions will be wrong.
- 标签绘制:
cv2.putTextoverlays the prediction on the frame. The(0,255,0)value is green; change it to(0,0,255)for red if you prefer. - 资源清理:Always call
cap.release()andcv2.destroyAllWindows()to free up camera resources and close OpenCV windows properly.
常见问题排查
- Camera won't open:Check if another app is using the camera, or verify your device's camera permissions.
- Inaccurate predictions:Double-check that preprocessing matches your training pipeline (e.g., same image size, normalization method, color space).
- Laggy feed:Try reducing the camera resolution (use
cap.set(cv2.CAP_PROP_FRAME_WIDTH, 640)andcap.set(cv2.CAP_PROP_FRAME_HEIGHT, 480)to lower resolution) or simplify preprocessing steps if possible.
内容的提问来源于stack exchange,提问作者Raju Sarkar
相关产品推荐
相关产品推荐

