如何从BeautifulSoup Tag对象中剥离外层标签?
实现方法
你可以通过以下两种常用方案实现需求:
方案1:直接提取标签内部内容(不修改原BeautifulSoup结构)
调用Tag对象的encode_contents()方法即可获取标签内部的完整内容(包含子标签结构),解码后就是目标结果:
from bs4 import BeautifulSoup # 示例初始化代码 raw_html = "<p>This is an example of something that I'm <strong>confused</strong> about</p>" soup = BeautifulSoup(raw_html, "html.parser") p_tag = soup.find("p") # 核心处理逻辑 target_content = p_tag.encode_contents().decode("utf-8") print(target_content)
输出结果:
This is an example of something that I'm <strong>confused</strong> about
方案2:使用unwrap()方法(修改原BeautifulSoup DOM结构)
unwrap()是原地修改方法,会直接移除调用该方法的标签、保留其内部内容到父节点中,不要直接取unwrap()的返回值,取修改后的根节点内容即可:
from bs4 import BeautifulSoup raw_html = "<p>This is an example of something that I'm <strong>confused</strong> about</p>" soup = BeautifulSoup(raw_html, "html.parser") p_tag = soup.find("p") # 核心处理逻辑 p_tag.unwrap() target_content = str(soup) print(target_content)
内容的提问来源于stack exchange,提问作者gammapoint
相关产品推荐
相关产品推荐

