如何提取Bransje对应dt标签后至下一个dt标签前的dd文本?
问题:提取指定dt标签后的连续dd文本(截止到下一个dt前)
需求是获取包含“Bransje”的dt标签之后、直至下一个dt标签前的所有dd标签内的文本,当前代码会把后续所有dd都输出,包括下一个dt后的内容,需要修正。
当前代码
from bs4 import BeautifulSoup html = """ <section class="panel"> <dl class="definition-list definition-list--inline"> <dt>Sektor</dt> <dd>Privat</dd> <dt>Sted</dt> <dd>Bratsbergveien 5, 7037 Trondheim</dd> <dt>Bransje</dt> <dd>Industri og produksjon,</dd> <dd>Maritim og offshore,</dd> <dd>Olje og gass</dd> <dt>Stillingsfunksjon</dt> <dd>Ingeniør</dd> </dl> </section> """ soup = BeautifulSoup(html,"lxml") for i in soup.select("dt:-soup-contains('Bransje') ~ dd"): print(i.text)
当前输出
Industri og produksjon, Maritim og offshore, Olje og gass Ingeniør
预期输出
Industri og produksjon, Maritim og offshore, Olje og gass
解决方案
原来的CSS选择器dt:-soup-contains('Bransje') ~ dd会选中该dt之后所有的dd标签,无法区分下一个dt的边界,因此需要换用遍历节点的方式来控制收集范围:
方法一:遍历目标dt的兄弟节点
找到目标dt后,逐个遍历其后续兄弟节点,遇到dd就输出,遇到新的dt就停止:
from bs4 import BeautifulSoup html = """ <section class="panel"> <dl class="definition-list definition-list--inline"> <dt>Sektor</dt> <dd>Privat</dd> <dt>Sted</dt> <dd>Bratsbergveien 5, 7037 Trondheim</dd> <dt>Bransje</dt> <dd>Industri og produksjon,</dd> <dd>Maritim og offshore,</dd> <dd>Olje og gass</dd> <dt>Stillingsfunksjon</dt> <dd>Ingeniør</dd> </dl> </section> """ soup = BeautifulSoup(html, "lxml") # 定位目标dt标签 target_dt = soup.find("dt", string="Bransje") # 从目标dt的下一个兄弟节点开始遍历 next_node = target_dt.next_sibling while next_node: # 跳过空白节点 if next_node.strip(): if next_node.name == "dd": print(next_node.text.strip()) # 遇到下一个dt就终止循环 elif next_node.name == "dt": break next_node = next_node.next_sibling
方法二:遍历dl的子节点
直接遍历dl的所有子节点,标记收集状态,遇到目标dt后开始收集dd,遇到新dt则停止:
from bs4 import BeautifulSoup html = """ <section class="panel"> <dl class="definition-list definition-list--inline"> <dt>Sektor</dt> <dd>Privat</dd> <dt>Sted</dt> <dd>Bratsbergveien 5, 7037 Trondheim</dd> <dt>Bransje</dt> <dd>Industri og produksjon,</dd> <dd>Maritim og offshore,</dd> <dd>Olje og gass</dd> <dt>Stillingsfunksjon</dt> <dd>Ingeniør</dd> </dl> </section> """ soup = BeautifulSoup(html, "lxml") dl = soup.find("dl", class_="definition-list definition-list--inline") collect = False for child in dl.children: # 跳过空白节点 if not child.strip(): continue if child.name == "dt": if child.string == "Bransje": collect = True elif collect: # 已经在收集状态,遇到新dt就停止 break elif collect and child.name == "dd": print(child.text.strip())
两种方法都能精准获取目标范围内的dd文本,避免包含后续dt下的内容。
内容的提问来源于stack exchange,提问作者SMTH
相关产品推荐
相关产品推荐

