解析含多命名空间的XML文件,提取CodeList节点并单独存文件
Got it, handling multi-namespace XML (especially CDISC ODM like your snippet) can be a bit tricky, but here are two practical approaches to get the job done—one using Python (great for flexibility) and another using XSLT (ideal if you prefer declarative stylesheets).
Option 1: Python with lxml Library
This is my go-to for complex XML tasks because lxml handles namespaces seamlessly and gives you fine-grained control.
Step 1: Install lxml
First, make sure you have the library installed:
pip install lxml
Step 2: Full Script
This script parses your XML, grabs every CodeList node, sanitizes the OID for safe filenames, and saves each node as a standalone XML file:
from lxml import etree # Map the default ODM namespace from your XML NAMESPACES = {'odm': 'http://www.cdisc.org/ns/odm/v1.3'} def extract_and_save_codelists(input_file_path): # Parse the original XML tree = etree.parse(input_file_path) # Find all CodeList nodes using XPath (namespace is critical here!) codelists = tree.xpath('//odm:CodeList', namespaces=NAMESPACES) for idx, codelist in enumerate(codelists): # Get the OID attribute to use as filename oid = codelist.get('OID') if not oid: print(f"Skipping CodeList #{idx+1} (no OID found)") continue # Sanitize OID to avoid invalid filename characters (adjust based on your OS) sanitized_oid = oid.replace('/', '_').replace('\\', '_').replace(':', '_').replace('?', '_') # Create a new XML tree with the CodeList as root new_tree = etree.ElementTree(codelist) # Write the file with proper formatting and XML declaration new_tree.write( f"{sanitized_oid}.xml", encoding='UTF-8', xml_declaration=True, pretty_print=True ) print(f"Saved: {sanitized_oid}.xml") # Replace with your actual input XML path extract_and_save_codelists("your_odm_file.xml")
Key Notes:
- Namespace Handling: The
NAMESPACESmap ensures we don't miss nodes—XPath ignores elements in default namespaces unless you explicitly reference them. - Filename Sanitization: The script replaces common invalid characters (like
/,:) to prevent file system errors. Tweak the replacement logic if you need to handle other edge cases. - Clean Output: Each file gets a proper UTF-8 declaration and pretty-printed formatting for readability.
Option 2: XSLT Stylesheet
If you prefer using XSLT (e.g., to run via command-line tools like xsltproc), here's a stylesheet that generates separate files using the EXSLT exsl:document extension:
<xsl:stylesheet version="1.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform" xmlns:odm="http://www.cdisc.org/ns/odm/v1.3" xmlns:exsl="http://exslt.org/common" extension-element-prefixes="exsl"> <xsl:output method="xml" encoding="UTF-8" indent="yes"/> <!-- Match the root and process all CodeList nodes --> <xsl:template match="/"> <xsl:apply-templates select="//odm:CodeList"/> </xsl:template> <!-- For each CodeList, save it to a file named after its OID --> <xsl:template match="odm:CodeList"> <exsl:document href="{@OID}.xml" method="xml" encoding="UTF-8" indent="yes"> <xsl:copy-of select="."/> </exsl:document> </xsl:template> </xsl:stylesheet>
Run It with xsltproc:
xsltproc your_stylesheet.xsl your_input_file.xml
Most modern XSLT processors (like Saxon or xsltproc) support the EXSLT extension needed for multiple output files.
内容的提问来源于stack exchange,提问作者ndesaintjorre

