如何使用Python将含多重复<topic>字段的复杂XML转换为CSV文件
Convert XML with Repeated
<topic> Fields to CSV Using Python Hey there, I've got you covered on this conversion. The key challenge here is merging those multiple <topic> elements into a single comma-separated cell while mapping all other fields correctly to CSV columns. Let's use Python's built-in libraries (no extra packages needed!) to get this done:
Step-by-Step Explanation & Code
First, we'll use xml.etree.ElementTree to parse the XML, and the csv module to write our output. Here's the full working code:
import xml.etree.ElementTree as ET import csv # 1. Parse your XML file (replace 'input.xml' with your actual file path) tree = ET.parse('input.xml') root = tree.getroot() # 2. Gather all unique field names to use as CSV headers csv_headers = set() for doc in root.findall('doc'): for field in doc.findall('field'): csv_headers.add(field.get('name')) # Convert to a sorted list for consistent column order (optional but recommended) csv_headers = sorted(csv_headers) # 3. Process each <doc> node to build CSV rows csv_rows = [] for doc in root.findall('doc'): row = {} topic_list = [] for field in doc.findall('field'): field_name = field.get('name') # Handle cases where field might have no text field_value = field.text.strip() if field.text else "" if field_name == 'topic': # Collect all topics for this doc topic_list.append(field_value) else: # Assign other fields directly to the row row[field_name] = field_value # Merge topics into a single comma-separated string row['topic'] = ', '.join(topic_list) csv_rows.append(row) # 4. Write the data to a CSV file with open('output.csv', 'w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=csv_headers) writer.writeheader() writer.writerows(csv_rows)
Key Details to Note:
- Handling Empty Fields: The code checks if
field.textexists before callingstrip()to avoid errors if a field has no content. - Header Consistency: Using a
setensures we only get unique field names (sotopiconly appears once in the header), and sorting gives a predictable column order. - Encoding: Using
encoding='utf-8'ensures special characters in your XML are preserved correctly in the CSV. - XML as String: If your XML is stored as a string instead of a file, replace
ET.parse('input.xml')withroot = ET.fromstring(your_xml_string).
If you run into edge cases (like duplicate non-topic fields in a single <doc>), let me know—but based on your sample XML, this should work perfectly.
内容的提问来源于stack exchange,提问作者joe marshal
相关产品推荐
相关产品推荐

