You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取Bransje对应dt标签后至下一个dt标签前的dd文本?

问题:提取指定dt标签后的连续dd文本(截止到下一个dt前)

需求是获取包含“Bransje”的dt标签之后、直至下一个dt标签前的所有dd标签内的文本,当前代码会把后续所有dd都输出,包括下一个dt后的内容,需要修正。

当前代码

from bs4 import BeautifulSoup

html = """
    <section class="panel">
        <dl class="definition-list definition-list--inline">
            <dt>Sektor</dt>
                <dd>Privat</dd>
            <dt>Sted</dt>
                <dd>Bratsbergveien 5, 7037 Trondheim</dd>
            <dt>Bransje</dt>
                <dd>Industri og produksjon,</dd>
                <dd>Maritim og offshore,</dd>
                <dd>Olje og gass</dd>
            <dt>Stillingsfunksjon</dt>
                <dd>Ingeniør</dd>
        </dl>
    </section>
"""
soup = BeautifulSoup(html,"lxml")
for i in soup.select("dt:-soup-contains('Bransje') ~ dd"):
    print(i.text)

当前输出

Industri og produksjon,
Maritim og offshore,
Olje og gass
Ingeniør

预期输出

Industri og produksjon,
Maritim og offshore,
Olje og gass

解决方案

原来的CSS选择器dt:-soup-contains('Bransje') ~ dd会选中该dt之后所有的dd标签,无法区分下一个dt的边界,因此需要换用遍历节点的方式来控制收集范围:

方法一:遍历目标dt的兄弟节点

找到目标dt后,逐个遍历其后续兄弟节点,遇到dd就输出,遇到新的dt就停止:

from bs4 import BeautifulSoup

html = """
    <section class="panel">
        <dl class="definition-list definition-list--inline">
            <dt>Sektor</dt>
                <dd>Privat</dd>
            <dt>Sted</dt>
                <dd>Bratsbergveien 5, 7037 Trondheim</dd>
            <dt>Bransje</dt>
                <dd>Industri og produksjon,</dd>
                <dd>Maritim og offshore,</dd>
                <dd>Olje og gass</dd>
            <dt>Stillingsfunksjon</dt>
                <dd>Ingeniør</dd>
        </dl>
    </section>
"""
soup = BeautifulSoup(html, "lxml")
# 定位目标dt标签
target_dt = soup.find("dt", string="Bransje")
# 从目标dt的下一个兄弟节点开始遍历
next_node = target_dt.next_sibling

while next_node:
    # 跳过空白节点
    if next_node.strip():
        if next_node.name == "dd":
            print(next_node.text.strip())
        # 遇到下一个dt就终止循环
        elif next_node.name == "dt":
            break
    next_node = next_node.next_sibling

方法二:遍历dl的子节点

直接遍历dl的所有子节点,标记收集状态,遇到目标dt后开始收集dd,遇到新dt则停止:

from bs4 import BeautifulSoup

html = """
    <section class="panel">
        <dl class="definition-list definition-list--inline">
            <dt>Sektor</dt>
                <dd>Privat</dd>
            <dt>Sted</dt>
                <dd>Bratsbergveien 5, 7037 Trondheim</dd>
            <dt>Bransje</dt>
                <dd>Industri og produksjon,</dd>
                <dd>Maritim og offshore,</dd>
                <dd>Olje og gass</dd>
            <dt>Stillingsfunksjon</dt>
                <dd>Ingeniør</dd>
        </dl>
    </section>
"""
soup = BeautifulSoup(html, "lxml")
dl = soup.find("dl", class_="definition-list definition-list--inline")
collect = False

for child in dl.children:
    # 跳过空白节点
    if not child.strip():
        continue
    if child.name == "dt":
        if child.string == "Bransje":
            collect = True
        elif collect:
            # 已经在收集状态,遇到新dt就停止
            break
    elif collect and child.name == "dd":
        print(child.text.strip())

两种方法都能精准获取目标范围内的dd文本,避免包含后续dt下的内容。

内容的提问来源于stack exchange,提问作者SMTH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 03:14:54