谷歌航班爬虫遇'NoneType'错误,求条件判断实现方案
解决BeautifulSoup爬取谷歌航班时的NoneType属性错误
问题核心是:当flight.find()找不到目标元素时会返回None,直接调用.text就会触发'NoneType' object has no attribute 'text'错误。解决逻辑很明确——先判断元素是否存在,再决定取值,不存在就赋值为None,让代码继续执行。
给你两种实用的处理写法:
写法1:分步判断(清晰直观,性能更优)
先获取元素对象,再检查是否为None,避免重复执行查找操作:
# 以常规价格字段为例 reg_price_elem = flight.find('div', class_='YMlIz FpEdX') Reg_price = reg_price_elem.text.strip() if reg_price_elem else None
写法2:一行式三元表达式(简洁)
把判断和取值合并成一行,适合简单场景:
Reg_price = flight.find('div', class_='YMlIz FpEdX').text.strip() if flight.find('div', class_='YMlIz FpEdX') else None
注意:这种写法会执行两次find(),如果爬取数据量较大,优先选写法1。
修改后的完整代码
把所有可能返回None的字段统一处理,避免后续报错:
from bs4 import BeautifulSoup import requests import time html_text = requests.get('https://www.google.com/travel/flights/search?tfs=CBwQAhoeagcIARIDSkZLEgoyMDIzLTA3LTAzcgcIARIDQU1TGh5qBwgBEgNBTVMSCjIwMjMtMDctMTNyBwgBEgNKRktwAYIBCwj___________8BQAFIAZgBAQ').text soup = BeautifulSoup(html_text, 'lxml') flights = soup.find_all('li', class_='pIav2d') for flight in flights: # 处理时间字段 czas_elem = flight.find('span', class_='mv1WYe') czas = czas_elem.text.strip() if czas_elem else None # 处理中转字段 stops_elem = flight.find('div', class_='EfT7Ae AdWm1c tPgKwe') stops = stops_elem.text.strip() if stops_elem else None # 处理低价字段 cheapP_elem = flight.find('div', class_='YMlIz FpEdX jLMuyc') cheapP = cheapP_elem.text.strip() if cheapP_elem else None # 处理常规价格字段 reg_price_elem = flight.find('div', class_='YMlIz FpEdX') Reg_price = reg_price_elem.text.strip() if reg_price_elem else None print(f''' Time: {czas} Stops: {stops} Cheapest: {cheapP} Regular Price: {Reg_price} ''')
额外提示:谷歌航班的页面结构经常更新,依赖class属性爬取容易失效,后续如果再次出现异常,优先检查页面元素的class或结构是否变化。
内容的提问来源于stack exchange,提问作者Zyrian
相关产品推荐
相关产品推荐

