You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量提取文件夹中Pascal VOC XML的指定<object>节点并保存?

批量处理Pascal VOC XML文件提取指定节点的自动化方案

实现思路

用Python标准XML解析库xml.etree.ElementTree批量遍历XML文件,保留原标注的基础结构(folder、filename、size等节点),仅筛选符合需求的<object>节点,最后导出新的XML文件。

步骤与代码

  1. 核心处理脚本
    以下脚本可直接运行,你只需根据自身需求修改筛选规则:
import os
import xml.etree.ElementTree as ET

def process_voc_xml(input_dir, output_dir, filter_condition):
    # 自动创建输出文件夹
    os.makedirs(output_dir, exist_ok=True)
    
    # 遍历输入文件夹中所有XML文件
    for filename in os.listdir(input_dir):
        if not filename.lower().endswith('.xml'):
            continue
        
        xml_path = os.path.join(input_dir, filename)
        # 解析原XML文件
        tree = ET.parse(xml_path)
        root = tree.getroot()
        
        # 创建新的<annotation>根节点
        new_annotation = ET.Element('annotation')
        
        # 复制原文件中除<object>外的所有节点(保留基础信息)
        for child in root:
            if child.tag != 'object':
                new_annotation.append(child)
        
        # 筛选符合条件的<object>节点并添加到新文件
        for obj in root.findall('object'):
            if filter_condition(obj):
                new_annotation.append(obj)
        
        # 保存新XML文件,指定UTF-8编码避免乱码
        new_tree = ET.ElementTree(new_annotation)
        output_path = os.path.join(output_dir, filename)
        new_tree.write(output_path, encoding='utf-8', xml_declaration=True)

# ---------------------- 自定义筛选规则 ----------------------
# 示例1:仅保留类别为"car"或"person"的目标
def filter_by_class(obj):
    class_name = obj.find('name').text.strip()
    return class_name in ['car', 'person']

# 示例2:仅保留非困难样本(difficult标签为0)
def filter_by_difficult(obj):
    difficult_tag = obj.find('difficult')
    if difficult_tag is None:
        return True  # 无difficult标签默认保留
    return difficult_tag.text.strip() == '0'

# 示例3:仅保留未被截断的目标(truncated标签为0)
def filter_by_truncated(obj):
    truncated_tag = obj.find('truncated')
    return truncated_tag is not None and truncated_tag.text.strip() == '0'
# ----------------------------------------------------------

if __name__ == '__main__':
    # 替换为你的实际路径
    INPUT_FOLDER = "./original_annotations"
    OUTPUT_FOLDER = "./filtered_annotations"
    
    # 选择你需要的筛选规则,比如用filter_by_class就传这个函数
    process_voc_xml(INPUT_FOLDER, OUTPUT_FOLDER, filter_by_class)

使用说明

  • 替换INPUT_FOLDER和OUTPUT_FOLDER为你的实际文件夹路径
  • 根据需求修改或新增筛选函数:比如需要筛选特定bndbox尺寸的目标,只需在筛选函数中获取obj.find('bndbox')下的xmin/ymin等节点做判断即可
  • 运行脚本后,输出文件夹中的XML文件会保留原文件的基础结构,仅包含符合规则的<object>节点,可直接用于TensorFlow自定义目标检测的数据集准备

内容的提问来源于stack exchange,提问作者Tahmid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 06:25:21