使用BeautifulSoup按部分类名提取HTML元素失效如何解决?
问题原因
BeautifulSoup的class_参数默认执行精确匹配,你需要匹配的元素class属性值为dErF-carrier-text,和代码中填写的carrier-text不完全相等,因此无法定位到目标元素。
解决方法
共有两种常用的模糊匹配方案:
方案1:CSS属性包含选择器(推荐)
使用select_one方法配合*=匹配规则,*=代表属性值包含指定片段即可,写法更简洁:
from bs4 import BeautifulSoup html = """ <div class="dErF-carrier-text"> Alaska Airlines 398 </div> """ soup = BeautifulSoup(html, 'html.parser') text = soup.select_one('div[class*="carrier-text"]').text # strip()用于去除内容前后多余的换行、空格 print(text.strip())
运行输出结果为:Alaska Airlines 398
方案2:正则表达式匹配
如果习惯使用find方法,可以配合正则表达式实现模糊匹配:
from bs4 import BeautifulSoup import re html = """ <div class="dErF-carrier-text"> Alaska Airlines 398 </div> """ soup = BeautifulSoup(html, 'html.parser') text = soup.find('div', class_=re.compile(r'carrier-text')).text print(text.strip())
- 若存在多个符合条件的元素,将
select_one替换为select、find替换为find_all即可获取全部匹配结果的列表。
内容的提问来源于stack exchange,提问作者Haotopia
相关产品推荐
相关产品推荐

