You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取HTML树中指定文本标签上方的内容?

提取指定标签上方的HTML内容

我有一棵HTML树,需要保留其中包含指定文本的某类标签上方的所有内容。示例中仅有一个文本为Notes的<b>标签,但实际场景中可能存在多个。

示例HTML

<br/>
Hello
<br/>
<b>
 Notes
</b>
<br/>
Hello
<a name="test">
  Hello2
</a>

期望处理结果

<br/>
Hello
<br/>

当前代码问题

我当前的代码仅能以列表形式得到期望结果,而非完整的HTML格式输出:

#book.html contains the example from above
openHtml = open('book.html', 'r')
soup = BeautifulSoup(openHtml, 'html.parser')
all=soup.find_all('b')
for i in all:
    if i.text.strip() == 'Notes':
        pos = all.index(i)
soup = soup.find_all("b")[pos].find_all_previous(string=True)
print(soup)

请问如何获取HTML格式的结果而非列表?


解决方案

要生成完整的HTML格式结果,你可以调整代码逻辑,直接提取并重组目标标签之前的节点:

from bs4 import BeautifulSoup

# 读取HTML文件
with open('book.html', 'r') as openHtml:
    soup = BeautifulSoup(openHtml, 'html.parser')

# 筛选出所有文本为"Notes"的<b>标签
target_tags = [tag for tag in soup.find_all('b') if tag.text.strip() == 'Notes']

# 处理每个目标标签(支持多个符合条件的标签场景)
for tag in target_tags:
    # 获取当前标签之前的所有节点,注意find_all_previous返回的是倒序节点
    prev_nodes = tag.find_all_previous()
    prev_nodes.reverse()
    
    # 创建新的BeautifulSoup对象来组装结果
    result_soup = BeautifulSoup("", 'html.parser')
    for node in prev_nodes:
        result_soup.append(node)
    
    # 输出格式化后的HTML内容
    print(result_soup.prettify())

说明

  • find_all_previous()会返回目标标签之前的所有节点,但顺序是从靠近目标标签的位置往前排,所以需要用reverse()恢复原文档顺序
  • 通过新建BeautifulSoup对象来拼接节点,最后用prettify()输出标准格式的HTML
  • 如果只需要处理第一个符合条件的标签,直接取target_tags[0]即可,无需循环

内容的提问来源于stack exchange,提问作者yemy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 06:50:24