使用BeautifulSoup如何提取不含子元素的标签且不修改原soup对象?
解决方案
方法1:直接构造同名同属性新标签(推荐,性能最优)
该方案不触碰原soup的任何节点,直接创建与目标div属性完全一致的新标签,天然没有子元素:
from bs4 import BeautifulSoup # 你的原始soup构造代码 html = """ <body> <div name='tag-i-want'> <span>I don't want this</span> </div> </body> """ soup = BeautifulSoup(html, 'html.parser') divs = [] for tag in soup.find_all('div'): # 直接创建同名称、同属性的新Tag对象 new_tag = BeautifulSoup.new_tag(soup, tag.name, attrs=tag.attrs) divs.append(str(new_tag)) print(divs)
输出结果:
["<div name='tag-i-want'></div>"]
方法2:深拷贝原标签后清空子节点
该方案适合需要保留标签的特殊元属性(如命名空间、自定义标记)的场景,深拷贝操作完全隔离原soup:
import copy from bs4 import BeautifulSoup # 同上构造原始soup html = """ <body> <div name='tag-i-want'> <span>I don't want this</span> </div> </body> """ soup = BeautifulSoup(html, 'html.parser') divs = [] for tag in soup.find_all('div'): # 深拷贝目标标签,与原soup完全隔离 temp_tag = copy.deepcopy(tag) # 清空所有子节点 temp_tag.clear() divs.append(str(temp_tag)) print(divs)
输出结果和方法1完全一致。
两种方案均不会修改原始soup对象,可放心使用,避免了
unwrap()方法修改原文档树的问题。
内容的提问来源于stack exchange,提问作者anon
相关产品推荐
相关产品推荐

