不使用etree从XML文件提取指定字符串的方法求助
没问题!既然没法用etree这类XML解析库,我们完全可以靠字符串处理或者正则表达式来提取目标内容,下面给你两个简单可行的方案:
方案1:字符串分割法(适合结构完全固定的XML片段)
如果你的XML片段格式非常规整(就像你给出的这样,标签紧凑、顺序固定),直接用字符串定位截取就足够简单高效:
xml_snippet = '<key>FDisplayName</key> <string>Dripo</string> <key>CFBundleIdentifier</key> <string>com.getdripo.dripo</string> <key>DTXcode</key>' # 先定位到CFBundleIdentifier对应的key标签结束的位置 start_pos = xml_snippet.find('<key>CFBundleIdentifier</key>') + len('<key>CFBundleIdentifier</key>') # 从该位置开始找后续<string>标签的起始点 string_open_tag_pos = xml_snippet.find('<string>', start_pos) + len('<string>') # 再找对应</string>标签的起始点 string_close_tag_pos = xml_snippet.find('</string>', string_open_tag_pos) # 截取中间的内容 bundle_id = xml_snippet[string_open_tag_pos:string_close_tag_pos] print(bundle_id) # 输出: com.getdripo.dripo
方案2:正则表达式法(更灵活,适应小幅度格式变化)
如果XML片段可能存在换行、多余空格这类格式变化,用正则表达式会更可靠,它能匹配不同格式下的目标内容:
import re xml_snippet = '<key>FDisplayName</key> <string>Dripo</string> <key>CFBundleIdentifier</key> <string>com.getdripo.dripo</string> <key>DTXcode</key>' # 正则规则:匹配CFBundleIdentifier的key标签,忽略中间任意空白,然后捕获<string>标签内的内容 pattern = r'<key>CFBundleIdentifier</key>\s*<string>([^<]+)</string>' match_result = re.search(pattern, xml_snippet) if match_result: bundle_id = match_result.group(1) print(bundle_id) # 输出: com.getdripo.dripo
这里的\s*会匹配任意数量的空格、换行符,([^<]+)则精准捕获<string>和</string>之间的所有内容(直到遇到下一个<为止),能应对大部分格式小变动。
这两个方案都不需要依赖任何XML解析工具,纯原生Python就能实现,完全符合你的场景需求。
内容的提问来源于stack exchange,提问作者CandyGum
相关产品推荐
相关产品推荐

