You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

不使用etree从XML文件提取指定字符串的方法求助

没问题!既然没法用etree这类XML解析库,我们完全可以靠字符串处理或者正则表达式来提取目标内容,下面给你两个简单可行的方案:

方案1:字符串分割法(适合结构完全固定的XML片段)

如果你的XML片段格式非常规整(就像你给出的这样,标签紧凑、顺序固定),直接用字符串定位截取就足够简单高效:

xml_snippet = '<key>FDisplayName</key> <string>Dripo</string> <key>CFBundleIdentifier</key> <string>com.getdripo.dripo</string> <key>DTXcode</key>'

# 先定位到CFBundleIdentifier对应的key标签结束的位置
start_pos = xml_snippet.find('<key>CFBundleIdentifier</key>') + len('<key>CFBundleIdentifier</key>')
# 从该位置开始找后续<string>标签的起始点
string_open_tag_pos = xml_snippet.find('<string>', start_pos) + len('<string>')
# 再找对应</string>标签的起始点
string_close_tag_pos = xml_snippet.find('</string>', string_open_tag_pos)
# 截取中间的内容
bundle_id = xml_snippet[string_open_tag_pos:string_close_tag_pos]
print(bundle_id)  # 输出: com.getdripo.dripo
方案2:正则表达式法(更灵活,适应小幅度格式变化)

如果XML片段可能存在换行、多余空格这类格式变化,用正则表达式会更可靠,它能匹配不同格式下的目标内容:

import re

xml_snippet = '<key>FDisplayName</key> <string>Dripo</string> <key>CFBundleIdentifier</key> <string>com.getdripo.dripo</string> <key>DTXcode</key>'

# 正则规则:匹配CFBundleIdentifier的key标签,忽略中间任意空白,然后捕获<string>标签内的内容
pattern = r'<key>CFBundleIdentifier</key>\s*<string>([^<]+)</string>'
match_result = re.search(pattern, xml_snippet)

if match_result:
    bundle_id = match_result.group(1)
    print(bundle_id)  # 输出: com.getdripo.dripo

这里的\s*会匹配任意数量的空格、换行符,([^<]+)则精准捕获<string>和</string>之间的所有内容(直到遇到下一个<为止),能应对大部分格式小变动。

这两个方案都不需要依赖任何XML解析工具,纯原生Python就能实现,完全符合你的场景需求。

内容的提问来源于stack exchange,提问作者CandyGum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:25:41