You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中startswith匹配含.的URL前缀失效问题求助

解决XML标签名前缀匹配失效的问题

我之前也碰到过类似的问题,其实是ElementTree处理XML命名空间的小坑在搞鬼!你遇到的情况是:用startswith匹配带点的命名空间前缀(比如"{http://www.d")或者特定域名的部分前缀时完全失效,只有当前缀写到}字符时才能正常匹配。

先看看你原来的代码:

#!/usr/bin/python
import xml.etree.ElementTree as ET
import re
tree = ET.parse('COMMON_08.xml')
root = tree.getroot()
searchString = "{http://www.d"
def recurse(node):
    for child in node:
        if child.tag.startswith(searchString):
            print(child.tag)
            #print(child.attrib)
            #print(child.text)
            recurse(child)
node = root.findall(".")
recurse(node)

问题根源

ElementTree在解析XML时,如果文档里有xmlns命名空间声明,会自动把长格式的{http://xxx...}tagname转换成前缀:tagname的简写形式(比如把{http://www.duolog.com/2011/05/socrates}property转成duolog:property)。这时候你用原始的命名空间前缀字符串(比如"{http://www.d")去匹配,自然和转换后的duolog:property对不上,导致匹配失效。只有当你匹配到}时,刚好和未被简写的完整命名空间前缀格式吻合,所以才能正常工作。

两种解决方案

方案一:用命名空间映射精准匹配

给目标命名空间定义别名,然后构造正确的匹配前缀,这样就能和child.tag的实际格式对齐:

#!/usr/bin/python
import xml.etree.ElementTree as ET

tree = ET.parse('COMMON_08.xml')
root = tree.getroot()

# 定义命名空间映射,别名可以随便取,对应完整的命名空间URL
namespaces = {
    'duolog': 'http://www.duolog.com/2011/05/socrates',
    'spirit': 'http://www.spiritconsortium.org/XMLSchema/SPIRIT/1685-2009'
}

def recurse(node):
    for child in node:
        # 匹配duolog命名空间下以"p"开头的标签(比如property)
        duolog_prefix = f'{{{namespaces["duolog"]}}}'
        if child.tag.startswith(duolog_prefix + 'p'):
            print(child.tag)
            recurse(child)
        # 匹配spirit命名空间下以"f"开头的标签
        spirit_prefix = f'{{{namespaces["spirit"]}}}'
        if child.tag.startswith(spirit_prefix + 'f'):
            print(child.tag)
            recurse(child)

# 直接递归根节点即可,不用findall(".")
recurse(root)

这里用f'{{{namespace_url}}}'是因为{在f-string里需要转义成{{,这样就能生成和child.tag一致的{http://xxx...}格式前缀。

方案二:用正则拆分命名空间和标签名

如果不想维护命名空间映射,可以用正则把标签拆成命名空间和标签名两部分,分别判断:

#!/usr/bin/python
import xml.etree.ElementTree as ET
import re

tree = ET.parse('COMMON_08.xml')
root = tree.getroot()

# 正则表达式匹配{命名空间}标签名的格式
tag_regex = re.compile(r'^\{(.*)\}(.*)$')

def recurse(node):
    for child in node:
        match_result = tag_regex.match(child.tag)
        if match_result:
            namespace = match_result.group(1)
            tag_name = match_result.group(2)
            # 判断命名空间是否以"http://www.d"开头,或者spirit命名空间下标签名以"f"开头
            if namespace.startswith('http://www.d') or (namespace == 'http://www.spiritconsortium.org/XMLSchema/SPIRIT/1685-2009' and tag_name.startswith('f')):
                print(child.tag)
                recurse(child)

recurse(root)

这个方法先拆分标签的结构,再分别对命名空间和标签名做判断,完全避开了ElementTree自动转换命名空间的问题。

内容的提问来源于stack exchange,提问作者user3521633

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 13:02:49