使用BeautifulSoup与Chrome Driver无法获取Kayak航班搜索的指定类名div标签
问题:无法通过类名识别航班搜索结果,Selenium抓取失败
我想通过类名识别有效的航班搜索结果列表,进而遍历抓取价格,但代码始终无法识别目标类名。我知道网站用JavaScript渲染页面,但觉得Selenium在页面渲染完成后应该能识别对应标签,请问问题出在哪?

相关代码
import time import subprocess import selenium from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from bs4 import BeautifulSoup import pandas as pd #from playsound import playsound import datetime import threading service = Service(executable_path='xxx') option = webdriver.ChromeOptions() option.add_argument("--headless=new") option.add_argument('--ignore-certificate-errors') option.add_argument("--no-sandbox") option.add_argument('disable-notifications') driver = webdriver.Chrome(service=service,options=option) def search(dep,arr,date): print(f'''Input: Date:{date},Departure: {dep} - Arrival: {arr}''') temp_url = 'https://www.kayak.com/flights/' base_url = temp_url + dep+'-'+arr+'/'+str(date)+'?sort=price_a&fs=stops=0' df_record = pd.DataFrame(columns=['deptime','arrtime','dep','arr', 'airline' ,'flightNum','price','ling']) print("before webdriver.ChromeOptions()") my_url = base_url driver.get(my_url) print(my_url) time.sleep(3) # set the time to wait till web fully loaded # wait for the close button to be visible and click it try: close_button = driver.find_element(By.XPATH, '//*[@class="nrc6"]') close_button.click() except: print("close is not found.") elem = driver.find_element("xpath","//*") source_code = elem.get_attribute("outerHTML") #print(source_code) bs = BeautifulSoup(source_code, 'html.parser') #print(bs) #expand drawing_url = bs.find_all('button', class_='nrc6') print(len(drawing_url)) # this shouldn't be zero if len(drawing_url)==0: return else: print(base_url)
问题排查与解决建议
静态等待不可靠
固定time.sleep(3)无法适配动态页面的加载速度,Kayak这类网站的航班数据加载耗时不稳定,3秒可能不足以完成渲染。改用Selenium的显式等待,确保目标元素出现后再执行操作:from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 替换time.sleep(3),最多等待10秒直到目标元素出现 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "nrc6")) )Headless模式被反爬检测
Kayak有反爬机制,--headless=new模式下浏览器特征容易被识别,导致页面无法正常渲染完整内容。可以关闭headless模式,或者添加模拟真实浏览器的参数:option.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") option.add_argument("--window-size=1920,1080")元素定位类名错误
你用的nrc6类名可能仅对应弹窗关闭按钮,而非航班结果的类名。打开Chrome开发者工具,直接定位航班结果的父元素,复制其真实类名或XPath(比如可能是resultInner这类专属类名)。无需混合BeautifulSoup
Selenium本身可以直接定位动态渲染的元素,没必要把页面源码传给BeautifulSoup解析,避免中间步骤丢失动态内容。直接用Selenium API遍历结果:# 替换成航班结果的真实类名 flight_results = driver.find_elements(By.CLASS_NAME, "航班结果类名") for result in flight_results: # 替换成价格元素的真实类名 price = result.find_element(By.CLASS_NAME, "价格类名").text print(price)触发反爬机制
Kayak会检测异常请求,可添加随机等待时间、模拟页面滚动等用户行为:# 滚动页面到底部,模拟用户浏览 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(1)
内容的提问来源于stack exchange,提问作者Erik Johnsson
相关产品推荐
相关产品推荐

