You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解析含多命名空间的XML文件,提取CodeList节点并单独存文件

Extract CodeList Nodes from Multi-Namespace XML and Save by OID

Got it, handling multi-namespace XML (especially CDISC ODM like your snippet) can be a bit tricky, but here are two practical approaches to get the job done—one using Python (great for flexibility) and another using XSLT (ideal if you prefer declarative stylesheets).

Option 1: Python with lxml Library

This is my go-to for complex XML tasks because lxml handles namespaces seamlessly and gives you fine-grained control.

Step 1: Install lxml

First, make sure you have the library installed:

pip install lxml

Step 2: Full Script

This script parses your XML, grabs every CodeList node, sanitizes the OID for safe filenames, and saves each node as a standalone XML file:

from lxml import etree

# Map the default ODM namespace from your XML
NAMESPACES = {'odm': 'http://www.cdisc.org/ns/odm/v1.3'}

def extract_and_save_codelists(input_file_path):
    # Parse the original XML
    tree = etree.parse(input_file_path)
    
    # Find all CodeList nodes using XPath (namespace is critical here!)
    codelists = tree.xpath('//odm:CodeList', namespaces=NAMESPACES)
    
    for idx, codelist in enumerate(codelists):
        # Get the OID attribute to use as filename
        oid = codelist.get('OID')
        if not oid:
            print(f"Skipping CodeList #{idx+1} (no OID found)")
            continue
        
        # Sanitize OID to avoid invalid filename characters (adjust based on your OS)
        sanitized_oid = oid.replace('/', '_').replace('\\', '_').replace(':', '_').replace('?', '_')
        
        # Create a new XML tree with the CodeList as root
        new_tree = etree.ElementTree(codelist)
        
        # Write the file with proper formatting and XML declaration
        new_tree.write(
            f"{sanitized_oid}.xml",
            encoding='UTF-8',
            xml_declaration=True,
            pretty_print=True
        )
        print(f"Saved: {sanitized_oid}.xml")

# Replace with your actual input XML path
extract_and_save_codelists("your_odm_file.xml")

Key Notes:

  • Namespace Handling: The NAMESPACES map ensures we don't miss nodes—XPath ignores elements in default namespaces unless you explicitly reference them.
  • Filename Sanitization: The script replaces common invalid characters (like /, :) to prevent file system errors. Tweak the replacement logic if you need to handle other edge cases.
  • Clean Output: Each file gets a proper UTF-8 declaration and pretty-printed formatting for readability.

Option 2: XSLT Stylesheet

If you prefer using XSLT (e.g., to run via command-line tools like xsltproc), here's a stylesheet that generates separate files using the EXSLT exsl:document extension:

<xsl:stylesheet version="1.0" 
                xmlns:xsl="http://www.w3.org/1999/XSL/Transform"
                xmlns:odm="http://www.cdisc.org/ns/odm/v1.3"
                xmlns:exsl="http://exslt.org/common"
                extension-element-prefixes="exsl">

    <xsl:output method="xml" encoding="UTF-8" indent="yes"/>

    <!-- Match the root and process all CodeList nodes -->
    <xsl:template match="/">
        <xsl:apply-templates select="//odm:CodeList"/>
    </xsl:template>

    <!-- For each CodeList, save it to a file named after its OID -->
    <xsl:template match="odm:CodeList">
        <exsl:document href="{@OID}.xml" method="xml" encoding="UTF-8" indent="yes">
            <xsl:copy-of select="."/>
        </exsl:document>
    </xsl:template>

</xsl:stylesheet>

Run It with xsltproc:

xsltproc your_stylesheet.xsl your_input_file.xml

Most modern XSLT processors (like Saxon or xsltproc) support the EXSLT extension needed for multiple output files.

内容的提问来源于stack exchange,提问作者ndesaintjorre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:24:20