You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于TensorFlow Object Detection API:视频边界框坐标与时间戳日志需求

Solution: Log Bounding Boxes with Timestamps & Generate DataFrame

Got it, let's break down how to modify your code to meet your requirements. We'll add timestamp tracking, convert normalized coordinates to image pixels, log each detection to the terminal, and finally build a structured DataFrame with all the data.

First, make sure you have pandas installed (if not, run pip install pandas). Then, here's the updated code with explanations:

import cv2
import numpy as np
import tensorflow as tf
import pandas as pd  # 导入pandas用于生成DataFrame

# 假设你的detection_graph、category_index、cap、out已经提前初始化完成

# 初始化存储检测数据的列表和框编号计数器
box_records = []
current_box_id = 0

with detection_graph.as_default():
    with tf.Session(graph=detection_graph) as sess:
        while cap.isOpened():
            ret, image_np = cap.read()
            if ret == True:
                # 获取当前帧的时间戳(单位:毫秒,转成秒更直观)
                timestamp_ms = cap.get(cv2.CAP_PROP_POS_MSEC)
                timestamp_sec = timestamp_ms / 1000.0
                
                # 获取图像的高度和宽度,用于坐标转换
                image_height, image_width = image_np.shape[:2]

                image_np_expanded = np.expand_dims(image_np, axis=0)
                image_tensor = detection_graph.get_tensor_by_name('image_tensor:0')
                boxes = detection_graph.get_tensor_by_name('detection_boxes:0')
                scores = detection_graph.get_tensor_by_name('detection_scores:0')
                classes = detection_graph.get_tensor_by_name('detection_classes:0')
                num_detections = detection_graph.get_tensor_by_name('num_detections:0')
                
                (boxes, scores, classes, num_detections) = sess.run(
                    [boxes, scores, classes, num_detections],
                    feed_dict={image_tensor: image_np_expanded})
                
                # 遍历所有检测结果,过滤低置信度的框(这里用0.5作为阈值,可按需调整)
                for i in range(int(num_detections[0])):
                    score = scores[0][i]
                    if score > 0.5:  # 只保留置信度达标的检测结果
                        # 从归一化坐标转换为图像实际像素坐标
                        ymin_norm, xmin_norm, ymax_norm, xmax_norm = boxes[0][i]
                        ymin = ymin_norm * image_height
                        xmin = xmin_norm * image_width
                        ymax = ymax_norm * image_height
                        xmax = xmax_norm * image_width

                        # 输出到终端
                        print(f"框编号: {current_box_id}, 时间戳: {timestamp_sec:.2f}s, 坐标: [ymin={ymin:.2f}, xmin={xmin:.2f}, ymax={ymax:.2f}, xmax={xmax:.2f}]")

                        # 将数据添加到记录列表
                        box_records.append({
                            '框编号': current_box_id,
                            '时间戳': timestamp_sec,
                            'ymin': ymin,
                            'xmin': xmin,
                            'ymax': ymax,
                            'xmax': xmax
                        })

                        # 框编号自增
                        current_box_id += 1

                # 原有的可视化和保存逻辑保持不变
                vis_util.visualize_boxes_and_labels_on_image_array(
                    image_np,
                    np.squeeze(boxes),
                    np.squeeze(classes).astype(np.int32),
                    np.squeeze(scores),
                    category_index,
                    use_normalized_coordinates=True,
                    line_thickness=8)
                out.write(image_np)
                cv2.imshow('Output',image_np)
                
                if cv2.waitKey(1) & 0xFF == ord('q'):
                    cv2.destroyAllWindows()
                    break
            else:
                break
        
        # 释放资源
        cap.release()
        out.release()
        cv2.destroyAllWindows()

        # 生成最终的DataFrame
        detection_df = pd.DataFrame(box_records, columns=['框编号', '时间戳', 'ymin', 'xmin', 'ymax', 'xmax'])
        # 可选:保存结果到CSV文件
        # detection_df.to_csv('detection_results.csv', index=False)
        print("\n检测数据已生成DataFrame:")
        print(detection_df.head())

Key Changes Explained:

  • Timestamp Tracking: We use cap.get(cv2.CAP_PROP_POS_MSEC) to fetch the current frame's timestamp in milliseconds, then convert it to seconds for readability.
  • Coordinate Conversion: Multiply normalized coordinates (0-1 range) by the image's actual height/width to get pixel-level coordinates, which matches your requirement for image-based coordinates.
  • Terminal Logging: For each valid detection (above the confidence threshold), we print a clear line with the box ID, timestamp, and converted coordinates.
  • DataFrame Construction: We collect all detection data in a list of dictionaries, then convert it to a pandas DataFrame with exactly the columns you specified. You can uncomment the to_csv line to save results to a file for later analysis.
  • Confidence Filter: We added a check for score > 0.5 to avoid logging low-confidence detections—adjust this threshold based on how strict you want your results to be.

内容的提问来源于stack exchange,提问作者Eoin Murnaghan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:48:25