You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python为XML文件中的现有文本添加XML元素

解决XML前置纯文本节点转元素的问题

场景说明

需要将XML中父元素(如示例中的<B>)内的前置纯文本数值,包装为指定元素(如<Z>),输入输出示例如下:

输入XML:

<A>
    <B id="254">
        12.34
        <C>Lore</C>
        <D>9</D> 
    </B>
</A>

目标XML:

<A>
    <B id="254">
        <Z>12.34</Z>
        <C>Lore</C>
        <D>9</D> 
    </B>
</A>

方案1:使用XSLT处理

XSLT可精准匹配并转换XML节点,以下样式表会将<B>元素下的第一个非空白纯文本节点包装为<Z>:

<xsl:stylesheet version="1.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform">
    <!-- 原样复制所有节点和属性 -->
    <xsl:template match="@*|node()">
        <xsl:copy>
            <xsl:apply-templates select="@*|node()"/>
        </xsl:copy>
    </xsl:template>

    <!-- 匹配<B>元素,处理其子节点 -->
    <xsl:template match="B">
        <xsl:copy>
            <xsl:apply-templates select="@*"/>
            <!-- 遍历<B>的子节点 -->
            <xsl:for-each select="node()">
                <!-- 判断是否是第一个非空白文本节点 -->
                <xsl:if test="position()=1 and self::text() and normalize-space(.) != ''">
                    <Z>
                        <xsl:value-of select="normalize-space(.)"/>
                    </Z>
                </xsl:if>
                <!-- 其他节点原样输出(包括空白文本和子元素) -->
                <xsl:if test="not(position()=1 and self::text() and normalize-space(.) != '')">
                    <xsl:copy-of select="."/>
                </xsl:if>
            </xsl:for-each>
        </xsl:copy>
    </xsl:template>
</xsl:stylesheet>

方案2:使用Python lxml库处理

用代码实现的话,lxml库可灵活操作XML节点:

from lxml import etree

# 解析XML
xml_content = """
<A>
    <B id="254">
        12.34
        <C>Lore</C>
        <D>9</D> 
    </B>
</A>
"""
root = etree.fromstring(xml_content.encode('utf-8'))

# 遍历所有<B>元素
for b_element in root.xpath('//B'):
    # 遍历子节点,找到第一个非空白文本节点
    for node in b_element.iterchildren():
        if isinstance(node, etree._ElementUnicodeResult) and node.strip():
            # 创建<Z>元素
            z_element = etree.Element('Z')
            z_element.text = node.strip()
            # 替换原文本节点
            b_element.replace(node, z_element)
            break

# 输出处理后的XML
print(etree.tostring(root, encoding='utf-8', pretty_print=True).decode('utf-8'))

内容的提问来源于stack exchange,提问作者JME

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 09:23:15