按属性值首个单词对XML进行排序的技术问询
Got it, let's tackle this XML sorting task. You need to sort the <view> elements inside the <viewpoints> node based on the first word in their name attribute. Below are two reliable approaches using Python—ideal for handling your thousand-line XML file efficiently.
Approach 1: Using Python's Built-in xml.etree.ElementTree
This method uses Python's standard library, so you don't need to install any extra packages. Perfect for a quick, no-fuss solution.
import xml.etree.ElementTree as ET # Parse your XML file tree = ET.parse('your_input.xml') root = tree.getroot() # Target the <viewpoints> container viewpoints = root.find('viewpoints') # Grab all <view> elements inside it views = viewpoints.findall('view') # Define how we'll sort: grab the first word from the 'name' attribute def sort_key(view): name_value = view.get('name', '') # Split the name into words, handle empty names gracefully first_word = name_value.split()[0] if name_value.split() else '' return first_word # Sort the views using our custom key sorted_views = sorted(views, key=sort_key) # Clear the original views and re-add them in sorted order for view in views: viewpoints.remove(view) for view in sorted_views: viewpoints.append(view) # Write the sorted XML to a new file (always backup your original first!) tree.write('sorted_output.xml', encoding='UTF-8', xml_declaration=True)
Quick Notes:
- Backup first: Never overwrite your original XML without making a copy—mistakes happen!
- Namespaces?: If your XML uses namespaces, adjust the
find/findallcalls to include the namespace prefix (e.g.,root.find('{http://your-namespace-url}viewpoints')). - Empty names: The script handles
<view>elements with emptynameattributes by sorting them to the top (you can tweak this if needed).
Approach 2: Using lxml (For Cleaner Output & Better Performance)
If you want prettier formatted output or need to handle more complex XML structures, lxml is a great upgrade. First install it with pip install lxml.
from lxml import etree # Parse the XML file tree = etree.parse('your_input.xml') root = tree.getroot() # Locate the <viewpoints> and its <view> children viewpoints = root.find('viewpoints') views = viewpoints.findall('view') # Reuse our sorting key function def sort_key(view): name_value = view.get('name', '') first_word = name_value.split()[0] if name_value.split() else '' return first_word # Sort and reinsert the views sorted_views = sorted(views, key=sort_key) for view in views: viewpoints.remove(view) for view in sorted_views: viewpoints.append(view) # Write with pretty printing for readability tree.write('sorted_output.xml', encoding='UTF-8', xml_declaration=True, pretty_print=True)
Bonus with lxml:
The pretty_print=True argument automatically indents the output XML, making your thousand-line file way easier to scan after sorting.
内容的提问来源于stack exchange,提问作者Gangula

