You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Beautiful Soup处理嵌套<figure>标签并获取目标HTML内容?

解决嵌套
标签的清理问题

针对嵌套的figure标签结构,可通过Beautiful Soup的节点筛选与操作实现预期的清理效果,具体步骤如下:

方法一:筛选子节点拼接结果

from bs4 import BeautifulSoup

text = '<figure><figcaption></figcaption><figure><img alt="dissident1.jpg" class="image-inline" src="https://example.com/dissident1.jpg" title="dissident1.jpg"/><figcaption>A file photo of XXX in prison, provided by his family. Credit: XXX</figcaption></figure><strong>Autopsy demand</strong></figure>'

# 解析HTML
soup = BeautifulSoup(text, 'html.parser')

# 定位最外层figure标签
outer_figure = soup.find('figure')

# 过滤外层figure内的空figcaption节点,保留有效内容
valid_nodes = [
    node for node in outer_figure.contents 
    if not (node.name == 'figcaption' and not node.get_text(strip=True))
]

# 将有效节点转为HTML字符串
cleaned_html = ''.join(str(node) for node in valid_nodes)
print(cleaned_html)

方法二:删除空节点后unwrap外层标签

from bs4 import BeautifulSoup

text = '<figure><figcaption></figcaption><figure><img alt="dissident1.jpg" class="image-inline" src="https://example.com/dissident1.jpg" title="dissident1.jpg"/><figcaption>A file photo of XXX in prison, provided by his family. Credit: XXX</figcaption></figure><strong>Autopsy demand</strong></figure>'

soup = BeautifulSoup(text, 'html.parser')

outer_figure = soup.find('figure')
# 定位并删除空的figcaption
empty_caption = outer_figure.find('figcaption', string=lambda s: not s.strip())
if empty_caption:
    empty_caption.decompose()

# 移除外层figure标签,保留内部内容
outer_figure.unwrap()

cleaned_html = str(soup)
print(cleaned_html)

说明

两种方法核心逻辑一致:

  1. 定位最外层的figure标签(find('figure')会返回嵌套结构中的第一个外层标签)
  2. 清除外层内部的空figcaption节点
  3. 保留内层完整的figure和后续的strong标签,最终输出清理后的HTML内容

内容的提问来源于stack exchange,提问作者Isyan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 13:07:28