如何从包含指定文本<b>标签的<div>中提取文本?
解决方法:提取包含指定标签的
中的目标文本
问题分析
你的代码中,work_order_no.next_sibling获取到的是<br>标签元素,而非后续的目标文本,因此调用add.text无法得到预期内容。以下是几种可行的解决方案:
方案1:遍历标签的后续兄弟节点提取文本
通过next_siblings遍历所有后续节点,筛选出有效文本节点并清理空白:
from bs4 import BeautifulSoup raw_text = """<div style="margin: 0 0 15px 0;"> <b>Description</b><br> Dryer is not working. Please have someone take a look at it </div>""" soup = BeautifulSoup(raw_text, "html.parser") desc_b = soup.find("b", text="Description") target_text = "" for sibling in desc_b.next_siblings: if sibling.string: target_text += sibling.string.strip() print(target_text)
方案2:直接获取文本并过滤前缀
找到包含目标标签的div,提取全部文本后去除"Description"前缀和多余空白:
from bs4 import BeautifulSoup raw_text = """<div style="margin: 0 0 15px 0;"> <b>Description</b><br> Dryer is not working. Please have someone take a look at it </div>""" soup = BeautifulSoup(raw_text, "html.parser") target_div = soup.find("div", string=lambda text: text and "Description" in text) full_text = target_div.get_text(strip=True, separator=" ") target_text = full_text.replace("Description", "").strip() print(target_text)
方案3:通过标签的父节点提取排除指定标签的文本
找到标签的父div,提取其中除、
之外的文本内容:
from bs4 import BeautifulSoup raw_text = """<div style="margin: 0 0 15px 0;"> <b>Description</b><br> Dryer is not working. Please have someone take a look at it </div>""" soup = BeautifulSoup(raw_text, "html.parser") desc_b = soup.find("b", text="Description") target_div = desc_b.parent target_text = "".join([node.strip() for node in target_div.contents if node.name not in ["b", "br"]]).strip() print(target_text)
以上方案均可输出预期结果:
Dryer is not working. Please have someone take a look at it
内容的提问来源于stack exchange,提问作者Ben David
相关产品推荐
相关产品推荐

