如何为BeautifulSoup的find_all编写多条件自定义过滤函数筛选指定td标签
修正后的完整实现代码
import re from bs4 import BeautifulSoup def custom_filter(tag): # 仅保留td标签,排除th if tag.name != 'td': return False # 匹配id格式为td[数字]_[数字] tag_id = tag.get('id', '') if not re.fullmatch(r'td\d+_\d+', tag_id): return False # 过滤掉style含display:none的标签 tag_style = tag.get('style', '').replace(' ', '') if 'display:none' in tag_style: return False return True # html替换为实际的页面源码字符串 parsed = BeautifulSoup(html, 'html.parser') filtered_tags = parsed.find_all(custom_filter)
原有代码的问题说明
- 导入语句错误:bs4库导出的类名为
BeautifulSoup,而非BeautifulSoup4 - 标签类型判断逻辑错误:
tag.select('td')是查询当前标签的子级td元素,无法判断当前标签本身是否为td,直接读取tag.name属性即可完成判断 - style过滤逻辑错误:你写的负向预查正则
(?!display:none)只要字符串任意位置不匹配该内容就会返回True,完全无法过滤掉display:none的标签,直接判断style字符串中是否包含对应内容更简单可靠,同时兼容style为空、无style属性的场景 - id匹配逻辑不严谨:未加锚点的正则会命中所有包含
td[数字]_[数字]子串的id,比如test_td123_45也会被错误匹配,用re.fullmatch可以严格匹配整个id值
如果你提供的示例HTML中的th是实际需要筛选的标签,只需要把tag.name != 'td'修改为tag.name == 'th'即可。
内容的提问来源于stack exchange,提问作者slachet
相关产品推荐
相关产品推荐

