You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取XML中非嵌套的pos标签值?

提取BeautifulSoup解析XML时的非嵌套<pos>标签值

原始XML结构

<?xml version="1.0" encoding="UTF-8" ?>
<main_heading timestamp="20220113">
<details>
    <offer id="11" parent_id="12">
        <name>Alpha</name>
        <pos>697</pos>
        <kat_pis>
            <pos kat="2">112</pos>
        </kat_pis>
    </offer>
    <offer id="12" parent_id="31">
        <name>Beta</name>
        <pos>099</pos>
        <kat_pis>
            <pos kat="2">113</pos>
        </kat_pis>
    </offer>
</details>
</main_heading>

当前代码及问题

原代码会提取所有<pos>标签内容,包括嵌套在<kat_pis>里的:

soup = BeautifulSoup(file, 'xml')

pos = []
for i in (soup.find_all('pos')):
    pos.append(i.text)

得到结果:['697', '112', '099', '113'],但只需要直接属于<offer>下的非嵌套<pos>值,即['697', '099']。

解决方案

方案1:直接定位<offer>下的直接子<pos>标签

先遍历所有<offer>标签,再在每个<offer>内只找直接子节点的<pos>:

soup = BeautifulSoup(file, 'xml')

pos = []
for offer in soup.find_all('offer'):
    direct_pos = offer.find('pos', recursive=False)
    if direct_pos:
        pos.append(direct_pos.text)

方案2:过滤父标签非<kat_pis>的<pos>

遍历所有<pos>标签,判断其父标签是否不是<kat_pis>:

soup = BeautifulSoup(file, 'xml')

pos = []
for pos_tag in soup.find_all('pos'):
    if pos_tag.parent.name != 'kat_pis':
        pos.append(pos_tag.text)

方案3:用CSS选择器直接筛选

利用CSS选择器的直接子元素语法>,精准选中<offer>的直接子<pos>:

soup = BeautifulSoup(file, 'xml')

pos = [tag.text for tag in soup.select('offer > pos')]

以上三种方法都能得到期望的结果['697', '099']。

内容的提问来源于stack exchange,提问作者x89

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 03:30:46