使用Python csv模块移除group字段不在指定列表的CSV行并排序输出XML
Got it, let's break this down into a practical, efficient solution using Python's built-in modules. I'll cover filtering, sorting (optimized for long outputList), and XML output—all tailored to your needs.
Step-by-Step Solution
First, let's start with the full code example, then walk through each part to explain how it works:
import csv import xml.etree.ElementTree as ET from xml.dom import minidom # Your predefined group order list (scales to long lists) outputList = ["GroupB", "GroupA", "GroupC"] # -------------------------- # 1. Read & Filter CSV Rows # -------------------------- filtered_rows = [] with open("input.csv", "r", newline="", encoding="utf-8") as csv_file: # Use DictReader to access columns by name (requires CSV has headers) csv_reader = csv.DictReader(csv_file) for row in csv_reader: # Keep only rows where 'group' is in your target list if row["group"] in outputList: filtered_rows.append(row) # -------------------------- # 2. Sort by outputList Order # -------------------------- # Optimize for long outputList: Use a lookup dict (O(1) lookups vs O(n) index calls) group_position_map = {group: idx for idx, group in enumerate(outputList)} sorted_rows = sorted( filtered_rows, key=lambda row: group_position_map[row["group"]] ) # -------------------------- # 3. Generate Prettified XML # -------------------------- def prettify_xml(element): """Make XML output human-readable with indentation""" rough_xml = ET.tostring(element, "utf-8") parsed_xml = minidom.parseString(rough_xml) return parsed_xml.toprettyxml(indent=" ") # Create root XML element root = ET.Element("dataset") # Add each sorted row as an <item> with column-specific child elements for row in sorted_rows: item = ET.SubElement(root, "item") for column, value in row.items(): child = ET.SubElement(item, column) child.text = value # Write to file or print to console with open("output.xml", "w", encoding="utf-8") as xml_file: xml_file.write(prettify_xml(root)) # Optional: Print result to console print(prettify_xml(root))
Key Details Explained
- Filtering: Using
csv.DictReaderlets us reference columns by their header name (likerow["group"]), which is cleaner than using indexes. We only keep rows that match youroutputListgroups. - Sorting Optimization: For long
outputList, using a dictionary to map groups to their positions avoids the slowoutputList.index()method (which scans the entire list every time). This makes sorting scale efficiently even with hundreds/thousands of groups. - XML Generation: We build the XML structure programmatically, then use
minidomto add indentation—so the output isn't just a single messy line. Each CSV row becomes an<item>element with children for each column.
Quick Notes for Edge Cases
- If your CSV doesn't have a header row, pass
fieldnames=["col1", "group", ...]tocsv.DictReaderto define column names manually. - Ensure case consistency between
outputListvalues and your CSV'sgroupfield (e.g., "GroupA" vs "groupa" are treated as different values). - For extremely large CSV files, process rows one at a time (instead of storing all filtered rows in memory) to keep memory usage low—just filter, sort (using the position map), and write to XML in a single loop.
内容的提问来源于stack exchange,提问作者CircularRecursion
相关产品推荐
相关产品推荐

