You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:基于文本标注文件创建VOC格式XML标注文件,处理<head>文本遇阻

Got it, let's get your VOC-format XML annotations sorted out. First, let's finish the partial snippet you have to match the example structure you provided, then I'll share a reusable approach that avoids processing <head> tag content.

Completed VOC XML from Your Partial Snippet

Here's how to expand your existing fragment into a valid, full VOC annotation file:

<annotation verified="yes">
  <folder>images</folder>
  <filename>raccoon-1.jpg</filename>
  <path>/Users/datitran/Desktop/raccoon/images/raccoon-1.jpg</path>
  <source>
    <database>Unknown</database>
  </source>
  <!-- Optional: Add image metadata and object annotations (standard for VOC) -->
  <size>
    <width>600</width> <!-- Replace with actual image width -->
    <height>400</height> <!-- Replace with actual image height -->
    <depth>3</depth>
  </size>
  <segmented>0</segmented>
  <!-- Example object annotation (repeat for each labeled object) -->
  <object>
    <name>raccoon</name>
    <pose>Unspecified</pose>
    <truncated>0</truncated>
    <difficult>0</difficult>
    <bndbox>
      <xmin>100</xmin>
      <ymin>50</ymin>
      <xmax>500</xmax>
      <ymax>350</ymax>
    </bndbox>
  </object>
</annotation>

Reusable Script to Generate XML (Skipping <head> Tags)

Since you can't process text inside <head> tags, here's a Python script that reads your existing text annotations, strips out any <head> block content, and outputs valid VOC XML:

import os
from xml.etree.ElementTree import Element, SubElement, tostring
from xml.dom.minidom import parseString
from PIL import Image  # Use to get actual image dimensions

def create_voc_xml(image_path, database="Unknown", verified="yes", skip_head_content=True):
    # Extract filename from path
    image_filename = os.path.basename(image_path)
    
    # Root annotation element
    annotation = Element('annotation', {'verified': verified})
    
    # Core elements matching your example
    SubElement(annotation, 'folder').text = 'images'
    SubElement(annotation, 'filename').text = image_filename
    SubElement(annotation, 'path').text = image_path
    
    source = SubElement(annotation, 'source')
    SubElement(source, 'database').text = database
    
    # Get actual image size (avoids hardcoding)
    with Image.open(image_path) as img:
        width, height = img.size
        depth = len(img.getbands())
    
    size = SubElement(annotation, 'size')
    SubElement(size, 'width').text = str(width)
    SubElement(size, 'height').text = str(height)
    SubElement(size, 'depth').text = str(depth)
    
    SubElement(annotation, 'segmented').text = '0'
    
    # Process your input annotation text (skip <head> content)
    with open('your_input_annotations.txt', 'r') as f:
        content = f.read()
        
        # Remove <head> blocks if enabled
        if skip_head_content:
            start_idx = content.find('<head>')
            end_idx = content.find('</head>')
            while start_idx != -1 and end_idx != -1:
                content = content[:start_idx] + content[end_idx+7:]
                start_idx = content.find('<head>')
                end_idx = content.find('</head>')
        
        # Add your logic here to parse cleaned content into object annotations
        # Example: If your text has lines like "raccoon 100 50 500 350"
        for line in content.strip().split('\n'):
            if not line:
                continue
            parts = line.split()
            obj_name = parts[0]
            xmin, ymin, xmax, ymax = map(int, parts[1:5])
            
            obj_elem = SubElement(annotation, 'object')
            SubElement(obj_elem, 'name').text = obj_name
            SubElement(obj_elem, 'pose').text = 'Unspecified'
            SubElement(obj_elem, 'truncated').text = '0'
            SubElement(obj_elem, 'difficult').text = '0'
            
            bndbox = SubElement(obj_elem, 'bndbox')
            SubElement(bndbox, 'xmin').text = str(xmin)
            SubElement(bndbox, 'ymin').text = str(ymin)
            SubElement(bndbox, 'xmax').text = str(xmax)
            SubElement(bndbox, 'ymax').text = str(ymax)
    
    # Pretty-print XML for readability
    rough_xml = tostring(annotation, 'utf-8')
    formatted_xml = parseString(rough_xml).toprettyxml(indent="  ")
    return formatted_xml

# Example usage
output_xml = create_voc_xml(
    image_path='/Users/datitran/Desktop/raccoon/images/raccoon-1.jpg'
)

# Save to file
with open('raccoon-1.xml', 'w') as f:
    f.write(output_xml)

Key Tips:

  • <head> Handling: The script repeatedly removes all content between <head> and </head> tags—adjust this if your <head> tags have attributes or variations.
  • Image Dimensions: Uses PIL to auto-fetch image width/height/depth, so no need to hardcode values.
  • Object Parsing: The example assumes your input text has lines with class names and bounding boxes. Modify the parsing logic to match your actual annotation format (e.g., XML snippets, CSV lines, etc.).

内容的提问来源于stack exchange,提问作者Abraham Jose

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:38:29