You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Beautiful Soup提取深度嵌套的<p>标签文本内容

解决方案

使用BeautifulSoup实现

核心思路是先定位到所有目标p标签共同的顶层父容器top-panel,再递归查找该容器下所有层级的p标签,不受中间嵌套结构限制,代码示例如下:

from bs4 import BeautifulSoup

# 替换为你的实际HTML内容
html_content = """
<div class="top-panel">
  <div class="inside-panel-0">
    <h1 class="h1-title">Some Title</h1>
  </div>
  <div class="inside-panel-0">
    <div class="inside-panel-1">
      <p> I want to extract this copy</p>
    </div>
    <div class="inside-panel-1">
      <p>I want to extract this copy</p>
    </div>
  </div>
</div>
"""

soup = BeautifulSoup(html_content, 'html.parser')
# 定位顶层容器,查找其下所有p标签
target_p = soup.find('div', class_='top-panel').find_all('p')
# 也可以用css选择器写法:target_p = soup.select('div.top-panel p')

# 提取文本
result = [p.get_text(strip=True) for p in target_p]
print(result)

运行后输出结果为:

['I want to extract this copy', 'I want to extract this copy']

其他可选方案:lxml+xpath

如果你追求更高的解析效率,也可以用lxml库的xpath语法实现跨层级查找,写法更简洁:

from lxml import etree

html = etree.HTML(html_content)
# 匹配top-panel下所有层级的p标签文本
result = [text.strip() for text in html.xpath('//div[@class="top-panel"]//p/text()') if text.strip()]
print(result)

你之前的方案失效是因为仅匹配了固定层级的直接子p标签,上述两种方法均采用全层级递归匹配逻辑,不受p标签上层嵌套结构的影响,只要在指定的顶层容器范围内即可被识别。

内容的提问来源于stack exchange,提问作者Freddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 14:15:03