如何用Python提取XML重复标签内的指定元素?
提取XML中Codes标签下的唯一代码并格式化输出
直接结合你使用的BeautifulSoup + Selenium流程给出解决方案:
步骤1:获取并解析XML内容
先通过Selenium拿到目标XML页面的源码,再用BeautifulSoup的XML解析器处理:
from selenium import webdriver from bs4 import BeautifulSoup # 初始化Chrome驱动(确保驱动路径正确,其他浏览器同理) driver = webdriver.Chrome() driver.get("你的目标XML文件URL") # 获取XML源码后关闭浏览器 xml_content = driver.page_source driver.quit() # 用BeautifulSoup解析XML(必须指定'xml'解析器) soup = BeautifulSoup(xml_content, 'xml')
步骤2:提取代码并去重
利用Python集合自动去重的特性,遍历所有Codes标签提取有效内容:
# 用集合存储唯一代码 unique_codes = set() # 遍历所有Codes标签 for code_tag in soup.find_all('Codes'): # 提取文本并清理首尾空白 code = code_tag.get_text(strip=True) # 跳过空内容的无效标签 if code: unique_codes.add(code) # 转成逗号分隔的字符串(可选排序,按需调整) result = ','.join(sorted(unique_codes)) print(result)
关键注意事项
- 标签大小写:XML区分大小写,
find_all('Codes')中的标签名必须和实际XML结构完全匹配(比如实际是小写codes就改成对应写法)。 - 嵌套结构处理:如果Codes标签内部还有子标签(如
<Codes><Code>72000000</Code></Codes>),需要调整提取逻辑:code = code_tag.find('Code').get_text(strip=True) - 大文件优化:如果XML体积特别大,用Python内置的
xml.etree.ElementTree解析速度更快,替代示例:import xml.etree.ElementTree as ET from selenium import webdriver driver = webdriver.Chrome() driver.get("你的目标XML文件URL") xml_content = driver.page_source driver.quit() root = ET.fromstring(xml_content) unique_codes = set() # 用XPath定位所有Codes标签 for code_tag in root.findall('.//Codes'): code = code_tag.text.strip() if code_tag.text else '' if code: unique_codes.add(code) result = ','.join(sorted(unique_codes)) print(result)
内容的提问来源于stack exchange,提问作者NukeSkull
相关产品推荐
相关产品推荐

