You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup优雅提取嵌套<p>标签中的指定内容?

优雅处理Beautiful Soup中嵌套

标签的内容提取需求

我完全懂你的痛点——这种结构相似但细节又有差异的嵌套标签,直接用常规方法要么漏内容要么混进无关信息,确实头疼。针对你给出的场景,我们可以拆解成三个明确的目标来逐个解决:提取example类span的文本、清理掉样式标签后的p标签正文、收集所有referencequote的内容。

核心思路

我们可以遍历每个包含目标span的<p>标签,对每个p标签分别处理:

  • 提取example值:直接定位p标签内的.example span,取其文本内容
  • 清理正文文本:复制当前p标签的副本,移除其中的.example和.referencequote元素,再用get_text()获取干净的正文(这样就能自动忽略<strong>、<em>这类样式标签了)
  • 收集引用内容:提取当前p标签内所有.referencequote的文本,存入列表

代码实现

from bs4 import BeautifulSoup

html_content = """
<p><span class="example" data-location="1:20">20</span>normal string</p> 
<p><span class="example" data-location="1:21">21</span>this text <strong>belongs together</strong></p> 
<p><span class="example" data-location="1:22">22</span>some text (<span class="referencequote">a reference text</span>)that might continue</p> 
<p><span class="example" data-location="1:23">23</span>more text</p><div class="linebreak"></div> 
<p><span class="example" data-location="1:22">24</span>text with (<span class="referencequote">first</span>)two references <span class="referencequote">second</span>.</p>
"""

soup = BeautifulSoup(html_content, "html.parser")
results = []

# 遍历所有包含.example span的p标签
for p_tag in soup.select("p:has(.example)"):
    # 提取example值
    example_text = p_tag.select_one(".example").get_text(strip=True)
    
    # 收集所有referencequote内容
    references = [ref.get_text(strip=True) for ref in p_tag.select(".referencequote")]
    
    # 创建p标签副本,移除不需要的元素来获取干净正文
    p_copy = p_tag.copy()
    # 移除example span和referencequote元素
    for elem in p_copy.select(".example, .referencequote"):
        elem.decompose()
    # 获取清理后的正文,处理多余空格
    clean_text = p_copy.get_text(strip=True).replace("  ", " ")
    
    results.append({
        "example": example_text,
        "text": clean_text,
        "reference": references
    })

# 打印验证结果
for item in results:
    print(f"example={item['example']}, text='{item['text']}', reference={item['reference']}")

代码说明

  • 使用p:has(.example)精准定位所有包含目标span的p标签,避免处理无关的p标签
  • 复制p标签再删除无关元素的方式,既保留了原标签结构用于提取样式化文本,又不会破坏原始soup结构
  • get_text()会自动合并<strong>这类标签内的文本,完美解决样式标签的问题
  • 最后用strip()和replace()处理多余的空格,让结果更整洁

运行这段代码后,就能得到你期望的精准结果啦!


内容的提问来源于stack exchange,提问作者phlprcks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 16:12:50