求助:基于文本标注文件创建VOC格式XML标注文件,处理<head>文本遇阻
Got it, let's get your VOC-format XML annotations sorted out. First, let's finish the partial snippet you have to match the example structure you provided, then I'll share a reusable approach that avoids processing <head> tag content.
Completed VOC XML from Your Partial Snippet
Here's how to expand your existing fragment into a valid, full VOC annotation file:
<annotation verified="yes"> <folder>images</folder> <filename>raccoon-1.jpg</filename> <path>/Users/datitran/Desktop/raccoon/images/raccoon-1.jpg</path> <source> <database>Unknown</database> </source> <!-- Optional: Add image metadata and object annotations (standard for VOC) --> <size> <width>600</width> <!-- Replace with actual image width --> <height>400</height> <!-- Replace with actual image height --> <depth>3</depth> </size> <segmented>0</segmented> <!-- Example object annotation (repeat for each labeled object) --> <object> <name>raccoon</name> <pose>Unspecified</pose> <truncated>0</truncated> <difficult>0</difficult> <bndbox> <xmin>100</xmin> <ymin>50</ymin> <xmax>500</xmax> <ymax>350</ymax> </bndbox> </object> </annotation>
Reusable Script to Generate XML (Skipping <head> Tags)
Since you can't process text inside <head> tags, here's a Python script that reads your existing text annotations, strips out any <head> block content, and outputs valid VOC XML:
import os from xml.etree.ElementTree import Element, SubElement, tostring from xml.dom.minidom import parseString from PIL import Image # Use to get actual image dimensions def create_voc_xml(image_path, database="Unknown", verified="yes", skip_head_content=True): # Extract filename from path image_filename = os.path.basename(image_path) # Root annotation element annotation = Element('annotation', {'verified': verified}) # Core elements matching your example SubElement(annotation, 'folder').text = 'images' SubElement(annotation, 'filename').text = image_filename SubElement(annotation, 'path').text = image_path source = SubElement(annotation, 'source') SubElement(source, 'database').text = database # Get actual image size (avoids hardcoding) with Image.open(image_path) as img: width, height = img.size depth = len(img.getbands()) size = SubElement(annotation, 'size') SubElement(size, 'width').text = str(width) SubElement(size, 'height').text = str(height) SubElement(size, 'depth').text = str(depth) SubElement(annotation, 'segmented').text = '0' # Process your input annotation text (skip <head> content) with open('your_input_annotations.txt', 'r') as f: content = f.read() # Remove <head> blocks if enabled if skip_head_content: start_idx = content.find('<head>') end_idx = content.find('</head>') while start_idx != -1 and end_idx != -1: content = content[:start_idx] + content[end_idx+7:] start_idx = content.find('<head>') end_idx = content.find('</head>') # Add your logic here to parse cleaned content into object annotations # Example: If your text has lines like "raccoon 100 50 500 350" for line in content.strip().split('\n'): if not line: continue parts = line.split() obj_name = parts[0] xmin, ymin, xmax, ymax = map(int, parts[1:5]) obj_elem = SubElement(annotation, 'object') SubElement(obj_elem, 'name').text = obj_name SubElement(obj_elem, 'pose').text = 'Unspecified' SubElement(obj_elem, 'truncated').text = '0' SubElement(obj_elem, 'difficult').text = '0' bndbox = SubElement(obj_elem, 'bndbox') SubElement(bndbox, 'xmin').text = str(xmin) SubElement(bndbox, 'ymin').text = str(ymin) SubElement(bndbox, 'xmax').text = str(xmax) SubElement(bndbox, 'ymax').text = str(ymax) # Pretty-print XML for readability rough_xml = tostring(annotation, 'utf-8') formatted_xml = parseString(rough_xml).toprettyxml(indent=" ") return formatted_xml # Example usage output_xml = create_voc_xml( image_path='/Users/datitran/Desktop/raccoon/images/raccoon-1.jpg' ) # Save to file with open('raccoon-1.xml', 'w') as f: f.write(output_xml)
Key Tips:
<head>Handling: The script repeatedly removes all content between<head>and</head>tags—adjust this if your<head>tags have attributes or variations.- Image Dimensions: Uses
PILto auto-fetch image width/height/depth, so no need to hardcode values. - Object Parsing: The example assumes your input text has lines with class names and bounding boxes. Modify the parsing logic to match your actual annotation format (e.g., XML snippets, CSV lines, etc.).
内容的提问来源于stack exchange,提问作者Abraham Jose
相关产品推荐
相关产品推荐

