基于TensorFlow Object Detection API:视频边界框坐标与时间戳日志需求
Solution: Log Bounding Boxes with Timestamps & Generate DataFrame
Got it, let's break down how to modify your code to meet your requirements. We'll add timestamp tracking, convert normalized coordinates to image pixels, log each detection to the terminal, and finally build a structured DataFrame with all the data.
First, make sure you have pandas installed (if not, run pip install pandas). Then, here's the updated code with explanations:
import cv2 import numpy as np import tensorflow as tf import pandas as pd # 导入pandas用于生成DataFrame # 假设你的detection_graph、category_index、cap、out已经提前初始化完成 # 初始化存储检测数据的列表和框编号计数器 box_records = [] current_box_id = 0 with detection_graph.as_default(): with tf.Session(graph=detection_graph) as sess: while cap.isOpened(): ret, image_np = cap.read() if ret == True: # 获取当前帧的时间戳(单位:毫秒,转成秒更直观) timestamp_ms = cap.get(cv2.CAP_PROP_POS_MSEC) timestamp_sec = timestamp_ms / 1000.0 # 获取图像的高度和宽度,用于坐标转换 image_height, image_width = image_np.shape[:2] image_np_expanded = np.expand_dims(image_np, axis=0) image_tensor = detection_graph.get_tensor_by_name('image_tensor:0') boxes = detection_graph.get_tensor_by_name('detection_boxes:0') scores = detection_graph.get_tensor_by_name('detection_scores:0') classes = detection_graph.get_tensor_by_name('detection_classes:0') num_detections = detection_graph.get_tensor_by_name('num_detections:0') (boxes, scores, classes, num_detections) = sess.run( [boxes, scores, classes, num_detections], feed_dict={image_tensor: image_np_expanded}) # 遍历所有检测结果,过滤低置信度的框(这里用0.5作为阈值,可按需调整) for i in range(int(num_detections[0])): score = scores[0][i] if score > 0.5: # 只保留置信度达标的检测结果 # 从归一化坐标转换为图像实际像素坐标 ymin_norm, xmin_norm, ymax_norm, xmax_norm = boxes[0][i] ymin = ymin_norm * image_height xmin = xmin_norm * image_width ymax = ymax_norm * image_height xmax = xmax_norm * image_width # 输出到终端 print(f"框编号: {current_box_id}, 时间戳: {timestamp_sec:.2f}s, 坐标: [ymin={ymin:.2f}, xmin={xmin:.2f}, ymax={ymax:.2f}, xmax={xmax:.2f}]") # 将数据添加到记录列表 box_records.append({ '框编号': current_box_id, '时间戳': timestamp_sec, 'ymin': ymin, 'xmin': xmin, 'ymax': ymax, 'xmax': xmax }) # 框编号自增 current_box_id += 1 # 原有的可视化和保存逻辑保持不变 vis_util.visualize_boxes_and_labels_on_image_array( image_np, np.squeeze(boxes), np.squeeze(classes).astype(np.int32), np.squeeze(scores), category_index, use_normalized_coordinates=True, line_thickness=8) out.write(image_np) cv2.imshow('Output',image_np) if cv2.waitKey(1) & 0xFF == ord('q'): cv2.destroyAllWindows() break else: break # 释放资源 cap.release() out.release() cv2.destroyAllWindows() # 生成最终的DataFrame detection_df = pd.DataFrame(box_records, columns=['框编号', '时间戳', 'ymin', 'xmin', 'ymax', 'xmax']) # 可选:保存结果到CSV文件 # detection_df.to_csv('detection_results.csv', index=False) print("\n检测数据已生成DataFrame:") print(detection_df.head())
Key Changes Explained:
- Timestamp Tracking: We use
cap.get(cv2.CAP_PROP_POS_MSEC)to fetch the current frame's timestamp in milliseconds, then convert it to seconds for readability. - Coordinate Conversion: Multiply normalized coordinates (0-1 range) by the image's actual height/width to get pixel-level coordinates, which matches your requirement for image-based coordinates.
- Terminal Logging: For each valid detection (above the confidence threshold), we print a clear line with the box ID, timestamp, and converted coordinates.
- DataFrame Construction: We collect all detection data in a list of dictionaries, then convert it to a pandas DataFrame with exactly the columns you specified. You can uncomment the
to_csvline to save results to a file for later analysis. - Confidence Filter: We added a check for
score > 0.5to avoid logging low-confidence detections—adjust this threshold based on how strict you want your results to be.
内容的提问来源于stack exchange,提问作者Eoin Murnaghan
相关产品推荐
相关产品推荐

