使用BeautifulSoup批量删除指定范围的et_pb_row_inner类标签及内容
优化实现方案
你可以直接用正则匹配类名规则,一次性筛选所有符合要求的元素,避免循环多次查询DOM,同时增加容错处理避免元素不存在时报错,优化后的代码如下:
import re from bs4 import BeautifulSoup def madisonsymphony(html_content): soup = BeautifulSoup(html_content, 'html.parser') # 删除所有header标签 for h in soup.find_all('header'): try: h.extract() except: pass # 删除所有footer标签 for f in soup.find_all('footer'): try: f.extract() except: pass # 删除指定id的顶部导航 tophead = soup.find("div",{"id":"top-header"}) if tophead: tophead.extract() # 批量匹配类名符合编号区间的div # 正则匹配et_pb_row_inner_后缀为2-22的规则 class_pattern = re.compile(r'et_pb_row_inner et_pb_row_inner_([2-9]|1[0-9]|2[0-2])\b') for div in soup.find_all("div", class_=class_pattern): div.extract() text = soup.getText(separator=u' ') return text
优化说明
- 效率提升:原实现需要循环21次遍历DOM查询元素,优化后仅需1次DOM遍历即可筛选出所有目标元素,HTML结构越复杂性能优势越明显
- 容错增强:给
tophead增加存在性判断,避免页面不存在对应id元素时直接抛出异常 - 通用性更强:如果后续需要调整删除的编号范围,仅需修改正则规则中的数字区间即可,无需调整循环参数
内容的提问来源于stack exchange,提问作者imhans33
相关产品推荐
相关产品推荐

