You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从BeautifulSoup Tag对象中剥离外层标签?

实现方法

你可以通过以下两种常用方案实现需求:

方案1:直接提取标签内部内容(不修改原BeautifulSoup结构)

调用Tag对象的encode_contents()方法即可获取标签内部的完整内容(包含子标签结构),解码后就是目标结果:

from bs4 import BeautifulSoup

# 示例初始化代码
raw_html = "<p>This is an example of something that I'm <strong>confused</strong> about</p>"
soup = BeautifulSoup(raw_html, "html.parser")
p_tag = soup.find("p")

# 核心处理逻辑
target_content = p_tag.encode_contents().decode("utf-8")
print(target_content)

输出结果:

This is an example of something that I'm <strong>confused</strong> about

方案2:使用unwrap()方法(修改原BeautifulSoup DOM结构)

unwrap()是原地修改方法,会直接移除调用该方法的标签、保留其内部内容到父节点中,不要直接取unwrap()的返回值,取修改后的根节点内容即可:

from bs4 import BeautifulSoup

raw_html = "<p>This is an example of something that I'm <strong>confused</strong> about</p>"
soup = BeautifulSoup(raw_html, "html.parser")
p_tag = soup.find("p")

# 核心处理逻辑
p_tag.unwrap()
target_content = str(soup)
print(target_content)

内容的提问来源于stack exchange,提问作者gammapoint

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 00:15:02