如何使用Beautiful Soup查找不匹配指定正则表达式的class?
可行,两种实现方法如下
方法1:使用正则表达式的负向前瞻实现反向匹配
直接在正则里添加负向前瞻断言,匹配不包含目标字符串的class属性值:
import re from bs4 import BeautifulSoup # 编译匹配不含"test"的正则 no_test_regex = re.compile(r'^(?!.*test.*).*$') # 查找所有class不包含"test"的div元素 for div in soup.find_all("div", {"class": no_test_regex}): print(div.get_text())
正则说明:(?!.*test.*)是负向前瞻,确保整个class字符串里不存在"test"子串;^和$锚定字符串首尾,避免部分匹配的误差。
方法2:使用lambda函数自定义判断逻辑
如果需要更灵活的条件(比如同时排除多个关键词),可以用lambda函数直接在find_all里做判断:
import re from bs4 import BeautifulSoup test_pattern = re.compile(r'.*test.*') # 查找class存在且不匹配test_pattern的div元素 for div in soup.find_all("div", class_=lambda c: c and not test_pattern.match(c)): print(div.get_text())
注:这里用class_而非class,是因为class是Python的保留关键字,BeautifulSoup用class_替代接收class属性的筛选条件;lambda里先判断c存在,避免没有class属性的元素被误选。
内容的提问来源于stack exchange,提问作者mangarapaul
相关产品推荐
相关产品推荐

