You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python将含多重复<topic>字段的复杂XML转换为CSV文件

Convert XML with Repeated <topic> Fields to CSV Using Python

Hey there, I've got you covered on this conversion. The key challenge here is merging those multiple <topic> elements into a single comma-separated cell while mapping all other fields correctly to CSV columns. Let's use Python's built-in libraries (no extra packages needed!) to get this done:

Step-by-Step Explanation & Code

First, we'll use xml.etree.ElementTree to parse the XML, and the csv module to write our output. Here's the full working code:

import xml.etree.ElementTree as ET
import csv

# 1. Parse your XML file (replace 'input.xml' with your actual file path)
tree = ET.parse('input.xml')
root = tree.getroot()

# 2. Gather all unique field names to use as CSV headers
csv_headers = set()
for doc in root.findall('doc'):
    for field in doc.findall('field'):
        csv_headers.add(field.get('name'))
# Convert to a sorted list for consistent column order (optional but recommended)
csv_headers = sorted(csv_headers)

# 3. Process each <doc> node to build CSV rows
csv_rows = []
for doc in root.findall('doc'):
    row = {}
    topic_list = []
    
    for field in doc.findall('field'):
        field_name = field.get('name')
        # Handle cases where field might have no text
        field_value = field.text.strip() if field.text else ""
        
        if field_name == 'topic':
            # Collect all topics for this doc
            topic_list.append(field_value)
        else:
            # Assign other fields directly to the row
            row[field_name] = field_value
    
    # Merge topics into a single comma-separated string
    row['topic'] = ', '.join(topic_list)
    csv_rows.append(row)

# 4. Write the data to a CSV file
with open('output.csv', 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.DictWriter(csv_file, fieldnames=csv_headers)
    writer.writeheader()
    writer.writerows(csv_rows)

Key Details to Note:

  • Handling Empty Fields: The code checks if field.text exists before calling strip() to avoid errors if a field has no content.
  • Header Consistency: Using a set ensures we only get unique field names (so topic only appears once in the header), and sorting gives a predictable column order.
  • Encoding: Using encoding='utf-8' ensures special characters in your XML are preserved correctly in the CSV.
  • XML as String: If your XML is stored as a string instead of a file, replace ET.parse('input.xml') with root = ET.fromstring(your_xml_string).

If you run into edge cases (like duplicate non-topic fields in a single <doc>), let me know—but based on your sample XML, this should work perfectly.

内容的提问来源于stack exchange,提问作者joe marshal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 07:32:30