如何使用Python解析文件夹中的多个XML文件?
Alright, let's break down how to parse all those XML files in your folder using Python. I'll cover two reliable approaches, using your sample XML snippet as a reference:
Python解析文件夹中多个XML文件的方法
方法一:用Python内置的xml.etree.ElementTree
This is the go-to option if you don't want to install any extra libraries—it's built right into Python, perfect for handling standard XML files.
Here's how it works:
- Loop through your target folder and pick out all files with the
.xmlextension - Read and parse each XML file one by one
- Extract the specific fields you need (like case name, AustLII link, citation details)
Code example:
import os import xml.etree.ElementTree as ET # Replace this with the actual path to your XML folder xml_folder = "./your_xml_directory" # Iterate over all XML files in the folder for filename in os.listdir(xml_folder): if filename.endswith(".xml"): file_path = os.path.join(xml_folder, filename) try: # Parse the XML file tree = ET.parse(file_path) root = tree.getroot() # Pull out the case name case_name = root.find("name").text print(f"Case Name: {case_name}") # Grab the AustLII link austlii_link = root.find("AustLII").text print(f"AustLII Link: {austlii_link}") # Extract citation details citations_section = root.find("citations") if citations_section is not None: for citation in citations_section.findall("citation"): # Quick note: Your sample XML has a syntax error here—`<citation "id=c0">` should be `<citation id="c0">` cite_class = citation.find("class").text cited_case = citation.find("tocase").text print(f"Citation Type: {cite_class}, Cited Case: {cited_case}") print("---") except ET.ParseError as e: print(f"Failed to parse {filename}: {e}") except Exception as e: print(f"Error processing {filename}: {e}")
方法二:用lxml库(更强大的XML处理工具)
If you need to handle more complex XML structures or want to use XPath queries for easier extraction, lxml is the way to go. First, you'll need to install it:
pip install lxml
Code example:
import os from lxml import etree xml_folder = "./your_xml_directory" for filename in os.listdir(xml_folder): if filename.endswith(".xml"): file_path = os.path.join(xml_folder, filename) try: # Parse the XML file tree = etree.parse(file_path) # Use XPath to get the case name case_name = tree.xpath("//name/text()")[0] print(f"Case Name: {case_name}") # Extract the AustLII link with XPath austlii_link = tree.xpath("//AustLII/text()")[0] print(f"AustLII Link: {austlii_link}") # Grab all citation entries citations = tree.xpath("//citations/citation") for cite in citations: cite_class = cite.xpath("./class/text()")[0] cited_case = cite.xpath("./tocase/text()")[0] print(f"Citation Type: {cite_class}, Cited Case: {cited_case}") print("---") except etree.XMLSyntaxError as e: print(f"XML syntax error in {filename}: {e}") except IndexError: print(f"Missing expected fields in {filename}") except Exception as e: print(f"Error processing {filename}: {e}")
A couple of important notes:
- Fix XML syntax first: Your sample XML has an invalid tag
<citation "id=c0">—this should be<citation id="c0">(attribute values need to be wrapped in quotes, and the attribute name comes first). Invalid XML will break any parser. - Handle encoding if needed: If your XML files use a non-UTF-8 encoding (like GB2312), you can specify it when parsing. For ElementTree, add
parser=ET.XMLParser(encoding='your_encoding')to theparse()call; for lxml, useparser=etree.XMLParser(encoding='your_encoding').
内容的提问来源于stack exchange,提问作者marisa
相关产品推荐
相关产品推荐

