如何用Python从XML提取顶点的最小/最大X、Y值以生成COCO JSON?
Fixing XML to COCO Bounding Box Extraction
Hey there! Let's sort out your XML parsing issue and get those bounding box coordinates efficiently. First, let's break down why your original code wasn't working:
What Went Wrong With Your Original Code?
The main issues were:
- You reset
minx = 100000inside the loop over each vertex, so you never tracked the global minimum across all vertices. - The line
for v in vertex.tag:was unnecessary—you were looping over the characters of the string "Vertex" instead of using the vertex data itself.
Clean, Efficient Solution
We can use XPath queries with Python's xml.etree.ElementTree to directly target the Vertex nodes, avoiding messy nested loops. Here's a streamlined implementation tailored to your goal of converting to COCO-style JSON:
import os import xml.etree.ElementTree as ET def extract_coco_annotations(xml_file_path): # Parse the XML file tree = ET.parse(xml_file_path) root = tree.getroot() # Initialize lists to hold coordinates and labels all_regions = [] # Iterate over each Region in the XML for region in root.findall(".//Regions/Region"): # Get the class label (e.g., "Benign") class_label = region.find(".//Attribute").get("Value") # Extract all Vertex coordinates under this Region vertices = region.findall("./Vertices/Vertex") x_coords = [] y_coords = [] for vertex in vertices: x = int(vertex.get("X")) y = int(vertex.get("Y")) x_coords.append(x) y_coords.append(y) # Calculate bounding box values min_x = min(x_coords) max_x = max(x_coords) min_y = min(y_coords) max_y = max(y_coords) # Convert to COCO-style bbox: [x_min, y_min, width, height] coco_bbox = [min_x, min_y, max_x - min_x, max_y - min_y] all_regions.append({ "class_label": class_label, "bbox": coco_bbox, "raw_coords": { "min_x": min_x, "max_x": max_x, "min_y": min_y, "max_y": max_y } }) return all_regions # Example usage xml_path = "path/to/your/xml/directory" for filename in os.listdir(xml_path): if filename.endswith(".xml"): xml_file = os.path.join(xml_path, filename) print(f"Processing: {xml_file}") annotations = extract_coco_annotations(xml_file) for idx, ann in enumerate(annotations): print(f"Region {idx+1}: {ann}")
Key Improvements:
- XPath Simplification: Queries like
.//Regions/Regionand./Vertices/Vertexlet us directly target the nodes we need, eliminating unnecessary nested loops. - Global Coordinate Tracking: We collect all coordinates first, then compute min/max values in one go—this is more efficient and avoids resetting values per vertex.
- COCO Compatibility: We include the standard COCO bbox format (
[x_min, y_min, width, height]) since that's your end goal. - Multi-Region Support: The code handles multiple
Regionelements per XML file, which is common in annotation datasets.
Quick Note:
Make sure your XML is well-formed (your sample had a typo: </Region should be </Region>—fix that before parsing to avoid errors!).
内容的提问来源于stack exchange,提问作者Suvidha
相关产品推荐
相关产品推荐

