You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取Kayak航班数据存CSV仅返回最后一条链接结果问题

问题定位方向
  • 变量作用域错误:检查detail_flights列表是不是在循环爬取每个链接的内部定义。如果每次循环都重新初始化该列表,之前的内容会被完全覆盖,最终只会保留最后一个链接的爬取结果。把总列表的定义放到循环外部即可避免覆盖问题。
  • 数据追加逻辑错误:检查单页数据爬取完成后,是不是用=赋值操作更新总列表,而非append()/extend()方法追加内容。赋值操作会直接清空原有列表内容,仅保留最后一次的赋值结果。
  • 输出/写入时机错误:如果打印列表、写入CSV的代码放在了循环内部,每爬一个链接就会输出一次当前的列表内容,就会出现累加重复输出的情况。仅需把这部分逻辑移到所有链接爬取完成的循环外部,就能只输出最终完整列表。
  • 页面数据残留问题:检查爬取每个新链接前,有没有等当前页元素完全加载、有没有清空上一页的临时数据缓存。页面未加载完成就提取数据,可能会拿到上一页的残留数据或者当前页的空数据。
修复代码示例
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import csv

# 替换为你自己的3个Kayak航班链接
target_urls = [
    "https://www.kayak.cn/xxx1",
    "https://www.kayak.cn/xxx2",
    "https://www.kayak.cn/xxx3"
]

# 总数据列表必须定义在循环外部,避免被覆盖
all_flight_data = []
driver = webdriver.Chrome()

for url in target_urls:
    driver.get(url)
    # 显式等待页面核心元素加载完成,比time.sleep更稳定
    WebDriverWait(driver, 15).until(
        EC.presence_of_element_located(("css selector", "你要提取的航班信息核心元素的选择器"))
    )
    # 替换为你自己的页面元素提取逻辑
    current_flight = {
        "航线": driver.find_element("css selector", "航线元素选择器").text,
        "价格": driver.find_element("css selector", "价格元素选择器").text,
        "起降时间": driver.find_element("css selector", "时间元素选择器").text
    }
    # 追加当前链接的爬取结果到总列表,不要用=赋值
    all_flight_data.append(current_flight)

# 所有链接爬取完成后统一打印、写入CSV,不要放在循环内部
print(all_flight_data)
with open("kayak_flights.csv", "w", encoding="utf-8-sig", newline="") as f:
    writer = csv.DictWriter(f, fieldnames=all_flight_data[0].keys())
    writer.writeheader()
    writer.writerows(all_flight_data)

driver.quit()
额外注意事项
  • 如果爬取过程中使用了临时变量存储单页提取结果,每次循环开始前要清空临时变量,避免残留上一页的数据。
  • 写入CSV时使用utf-8-sig编码可以避免中文乱码问题。

内容的提问来源于stack exchange,提问作者Box

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 06:18:01