You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取table与指定div间数量不固定的p标签

BeautifulSoup 提取两节点间不定量p标签实现方案

核心思路是通过BeautifulSoup原生的兄弟节点遍历能力,以两个已知位置的节点为边界做定向收集,全程返回原生Tag对象,完全兼容后续soup解析操作。

实现步骤

  • 先定位两个节点的公共父容器,也就是class="foo"的div标签,在容器内分别找到两个边界节点:作为起始边界的<table>标签、作为终止边界的class="bar"div标签
  • 从table标签的下一个兄弟节点开始向后逐节点遍历,碰到终止边界的bar div时立刻停止遍历
  • 遍历过程中自动跳过标签间换行、缩进生成的空白文本节点,收集所有命中的p标签即可

示例代码

from bs4 import BeautifulSoup, Tag

# 定位公共父容器
foo_container = soup.find("div", class_="foo")
# 定位起止边界节点
start_table = foo_container.find("table")
end_bar_div = foo_container.find("div", class_="bar")

target_p_tags = []
current_node = start_table.next_sibling
# 遍历到终止节点即停止
while current_node is not None and current_node != end_bar_div:
    # 仅收集标签类型为p的节点,跳过空白文本节点
    if isinstance(current_node, Tag) and current_node.name == "p":
        target_p_tags.append(current_node)
    current_node = current_node.next_sibling

方案优势

  • 无固定索引依赖,两个边界间不管是1个还是5个甚至更多p标签都能精准采集
  • 采集结果均为标准BeautifulSoup Tag对象,可直接调用.get_text()、.find()、.select()等任意soup方法开展后续提取,不存在正则方案丢失对象操作能力的问题
  • 碰到终止节点立刻停止遍历,不会误采集bar div内部的p标签,结果范围完全符合预期

内容的提问来源于stack exchange,提问作者Aadhiraj Nayar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 11:27:21