如何用BeautifulSoup移除HTML标签并保留内容?替代LXML strip_tag方法
在BeautifulSoup中移除标签但保留内容的方法
BeautifulSoup里没有直接对应lxml strip_tag的方法,但可以用unwrap()方法实现完全相同的效果——移除指定标签,同时保留标签内的所有内容(包括子标签和文本)。
基础用法
针对单个标签:
from bs4 import BeautifulSoup html = '<div class="extra"><span>需要保留的内容</span></div>' soup = BeautifulSoup(html, 'html.parser') # 选中要处理的标签,调用unwrap() target_tag = soup.find('div') target_tag.unwrap() print(soup) # 输出: <span>需要保留的内容</span>
批量处理多个标签
如果要一次性移除所有<span>和<div>标签,保留内部内容,可以遍历所有匹配的标签并调用unwrap():
html = ''' <div class="wrapper"> <span class="red">文本1</span> <div class="blue"> <p>嵌套内容</p> </div> </div> ''' soup = BeautifulSoup(html, 'html.parser') # 遍历所有需要移除的标签 for tag in soup.find_all(['span', 'div']): tag.unwrap() print(soup.prettify()) # 输出: # 文本1 # # <p> # 嵌套内容 # </p>
注意事项
unwrap()会直接修改原BeautifulSoup对象,不需要重新赋值- 如果标签是文档的根节点,
unwrap()会报错,建议先检查标签是否有父节点 - 若只想移除特定属性的标签,可以在
find_all()里添加过滤条件,比如soup.find_all('div', class_='extra')
内容的提问来源于stack exchange,提问作者mrgou
相关产品推荐
相关产品推荐

