You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取XML重复标签内的指定元素?

提取XML中Codes标签下的唯一代码并格式化输出

直接结合你使用的BeautifulSoup + Selenium流程给出解决方案:

步骤1:获取并解析XML内容

先通过Selenium拿到目标XML页面的源码,再用BeautifulSoup的XML解析器处理:

from selenium import webdriver
from bs4 import BeautifulSoup

# 初始化Chrome驱动(确保驱动路径正确,其他浏览器同理)
driver = webdriver.Chrome()
driver.get("你的目标XML文件URL")

# 获取XML源码后关闭浏览器
xml_content = driver.page_source
driver.quit()

# 用BeautifulSoup解析XML(必须指定'xml'解析器)
soup = BeautifulSoup(xml_content, 'xml')

步骤2:提取代码并去重

利用Python集合自动去重的特性,遍历所有Codes标签提取有效内容:

# 用集合存储唯一代码
unique_codes = set()

# 遍历所有Codes标签
for code_tag in soup.find_all('Codes'):
    # 提取文本并清理首尾空白
    code = code_tag.get_text(strip=True)
    # 跳过空内容的无效标签
    if code:
        unique_codes.add(code)

# 转成逗号分隔的字符串(可选排序,按需调整)
result = ','.join(sorted(unique_codes))
print(result)

关键注意事项

  • 标签大小写:XML区分大小写,find_all('Codes')中的标签名必须和实际XML结构完全匹配(比如实际是小写codes就改成对应写法)。
  • 嵌套结构处理:如果Codes标签内部还有子标签(如<Codes><Code>72000000</Code></Codes>),需要调整提取逻辑:
    code = code_tag.find('Code').get_text(strip=True)
    
  • 大文件优化:如果XML体积特别大,用Python内置的xml.etree.ElementTree解析速度更快,替代示例:
    import xml.etree.ElementTree as ET
    from selenium import webdriver
    
    driver = webdriver.Chrome()
    driver.get("你的目标XML文件URL")
    xml_content = driver.page_source
    driver.quit()
    
    root = ET.fromstring(xml_content)
    unique_codes = set()
    
    # 用XPath定位所有Codes标签
    for code_tag in root.findall('.//Codes'):
        code = code_tag.text.strip() if code_tag.text else ''
        if code:
            unique_codes.add(code)
    
    result = ','.join(sorted(unique_codes))
    print(result)
    

内容的提问来源于stack exchange,提问作者NukeSkull

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 04:45:35