解决Kayak航班页面URL爬取的BrokenPipeError问题
问题描述
尝试爬取Kayak多段行程航班搜索页面中的所有航班卡片URL时,运行代码出现BrokenPipeError: [Errno 32] Broken pipe错误,请求获取正确的爬取代码。
原代码如下:
url = 'https://www.kayak.com/flights/AMS-WMI,nearby/2023-02-15/WMI-SOF,nearby/2023-02-18/SOF-BEG,nearby/2023-02-20/BEG-MIL,nearby/2023-02-23/MIL-AMS,nearby/2023-02-25/?sort=bestflight_a&fs=stops=-2&attempt=1&lastms=1675195877028' requests = 0 chrome_options = webdriver.ChromeOptions() agents = ["Firefox/66.0.3","Chrome/73.0.3683.68","Edge/16.16299"] print("User agent: " + agents[(requests%len(agents))]) chrome_options.add_argument('--user-agent=' + agents[(requests%len(agents))] + '"') chrome_options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome('/Users/junerodriguez/Downloads/chromedriver_mac_arm64/chromedriver') driver.implicitly_wait(10) driver.get(url) sleep(randint(8,10)) xp_hrefs = "//div[@class='above-button']//a[contains(@class,'booking-link')]/href[@class='col col-best']" hrefs = driver.find_elements_by_xpath(xp_hrefs) hrefs
航班卡片结构参考:
修复方案
错误根源
BrokenPipeError多由ChromeDriver与浏览器版本不匹配、反爬机制触发连接中断、参数配置错误导致。- 原XPath路径错误:错误地将
href当作子节点,实际应通过@href属性提取链接;定位路径也不准确。 - User-Agent参数多了一个引号,导致格式非法。
- 缺少反爬配置,容易被Kayak识别为爬虫。
正确代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import random import time url = 'https://www.kayak.com/flights/AMS-WMI,nearby/2023-02-15/WMI-SOF,nearby/2023-02-18/SOF-BEG,nearby/2023-02-20/BEG-MIL,nearby/2023-02-23/MIL-AMS,nearby/2023-02-25/?sort=bestflight_a&fs=stops=-2&attempt=1&lastms=1675195877028' chrome_options = webdriver.ChromeOptions() # 使用标准完整UA字符串,随机选择 agents = [ "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/73.0.3683.68 Safari/537.36", "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:66.0) Gecko/20100101 Firefox/66.0", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Edge/16.16299 Safari/537.36" ] selected_agent = random.choice(agents) print(f"User agent: {selected_agent}") chrome_options.add_argument(f'--user-agent={selected_agent}') # 增强反爬配置,规避自动化检测 chrome_options.add_experimental_option('useAutomationExtension', False) chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_argument("--start-maximized") # 初始化驱动(新版Selenium无需手动指定chromedriver,版本不匹配需更新) driver = webdriver.Chrome(options=chrome_options) wait = WebDriverWait(driver, 20) try: driver.get(url) # 处理Cookie授权弹窗(无弹窗则跳过) try: accept_cookie = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Accept')]"))) accept_cookie.click() time.sleep(random.randint(2, 3)) except: pass # 等待航班链接加载完成,修正XPath定位 flight_elements = wait.until(EC.presence_of_all_elements_located( (By.XPATH, "//div[contains(@class, 'above-button')]//a[contains(@class, 'booking-link')]") )) # 提取有效链接 flight_links = [elem.get_attribute('href') for elem in flight_elements if elem.get_attribute('href')] print(f"共抓取到 {len(flight_links)} 个航班链接:") for idx, link in enumerate(flight_links, 1): print(f"{idx}. {link}") finally: # 确保浏览器关闭,避免资源泄漏 driver.quit()
关键修复点
- 修正User-Agent参数的引号错误,使用完整标准UA字符串,避免被识别。
- 添加反爬配置:禁用自动化标识、排除自动化开关,模拟真实用户环境。
- 替换隐式等待为显式等待,确保元素加载完成后再操作,避免页面未就绪导致的错误。
- 修正XPath路径:正确提取
a标签的href属性,而非错误的子节点写法。 - 添加Cookie弹窗处理逻辑,避免弹窗遮挡元素。
- 使用
try-finally确保浏览器正常关闭,防止连接残留引发BrokenPipeError。
内容的提问来源于stack exchange,提问作者June Smith
相关产品推荐
相关产品推荐

