You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬虫脚本无报错但无法生成CSV输出求助

问题描述

使用Selenium脚本抓取giffgaff平台的翻新手机信息(如iPhone 12 5G、Galaxy S21+ 5G的翻新页面),脚本配置为每日执行一次。初期1-2天运行正常,现在无法生成CSV输出。已尝试更换IP、改用geckodriver,问题依旧,且终端无任何报错信息。附上原脚本:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait 
from selenium.webdriver.support import expected_conditions as EC
import csv

options1 = webdriver.ChromeOptions()
    # options1.add_argument('--start-maximized')
    options1.add_experimental_option('excludeSwitches', ['enable-logging'])
    # options1.add_argument("window-size=1920x1480")
    # options1.add_argument("disable-dev-shm-usage")
    # options1.add_argument("--headless")
    driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options1)
    

    filename = "data.csv"
    with open(filename, 'w', encoding='utf-8', newline='') as file:
        writer = csv.writer(file)
        writer.writerow([ "URL", "MODEL", "STORAGE", "COLOUR", "GRADE", "PRICE"])

    List_open = open("links.txt")
    read_list = List_open.read()
    line_in_list = read_list.split("\n")

    for url in line_in_list:
        page = driver.execute_script("window.open('{}','_blank');".format(url))
        handles = driver.window_handles
        main_window = driver.current_window_handle
        driver.switch_to.window(handles[-1])
        sleep(5)

        try:
            wait=WebDriverWait(driver,10)
            cookies = wait.until(EC.visibility_of_element_located((By.XPATH,"//div[@class='cbot-layout--2column cbot-layout--right cbot-layout--ctas']/a[1]/span[1]")))
            cookies.click()
        except:
            pass
        
        sleep(5)
        model_name = wait.until(EC.visibility_of_element_located((By.XPATH,"//h1[@class='gg-u-text-shout']"))).text
        try:
            storage_container = wait.until(EC.visibility_of_all_elements_located((By.XPATH,"//div[@class='MemorySelector__MemoryContainer-jkajbs-0 jKQdwX gg-u-mb-small']/fieldset/ol/li/label")))
            for item in storage_container:
                size = item.text
                item.click()
                sleep(2)
                colour_container = driver.find_elements("xpath","//fieldset[@data-testid='colour-selector']/ol/li/label/span")
                for c in colour_container:
                    colour = c.text
                    grade_price = driver.find_elements("xpath","//div[@class='gg-u-pb-medium']/fieldset/ol/li")
                    for g in grade_price:
                        grade = g.find_element("xpath", ".//label/span[1]").text
                        price = g.find_element("xpath", ".//label/span[2]/span").text
                        
                        # print(url + "," + model_name + "," + size + "," + colour + "," + grade + "," + price.strip() + "\n")

                        with open(filename, 'a', encoding='utf-8') as file:
                            file.write(url + "," + model_name + "," + size + "," + colour + "," + grade + "," + price.strip() + "\n")
        except:
            with open(filename, 'a', encoding='utf-8') as file:
                    file.write(url + "," + model_name + "," + "None" + "," + "None" + "," + "Out of Stock" + "," + "Out of Stock" + "\n")
      
        driver.close()
        driver.switch_to.window(main_window)
问题排查与修复方案

以下是针对脚本问题的逐一排查和修复建议:

1. 无差别异常捕获吞噬所有错误

原脚本使用except:捕获所有异常,会掩盖元素定位失败、文件写入错误、网络超时等问题,导致无法定位根源。

  • 修复:捕获具体异常类型(如TimeoutException、NoSuchElementException),添加错误日志输出,明确问题点。

2. 动态元素定位器失效

页面中的动态类名(如MemorySelector__MemoryContainer-jkajbs-0 jKQdwX)是前端框架生成的,网站更新后会失效,触发异常后被吞掉,导致后续逻辑不执行。

  • 修复:改用静态属性定位,比如data-testid或页面稳定文本:
    • Cookie按钮可替换为//button[contains(text(), 'Accept all cookies')];
    • 存储选择器改用//fieldset[@data-testid='memory-selector']。

3. 文件写入逻辑存在隐患

多次打开/关闭CSV文件可能导致句柄异常,手动拼接字符串写入还会出现格式错误(如内容含逗号时破坏CSV结构)。

  • 修复:使用csv.writer的writerow方法统一写入,保持文件句柄在循环外打开,减少IO操作。

4. 硬编码等待不可靠

大量sleep(5)这类固定等待时间,会因网络延迟或页面加载慢导致元素未就绪,触发异常。

  • 修复:统一使用WebDriverWait显式等待,仅在必要时使用短暂sleep等待页面更新。

5. 代码缩进错误

原脚本中options1的配置代码存在缩进错误,部分环境下会导致脚本执行失败,但被异常捕获吞掉后无提示。

  • 修复:修正代码缩进,确保所有语句块层级正确。

6. 反爬机制触发

网站可能检测到自动化脚本,限制了元素交互或返回空白页,但脚本未处理这类场景。

  • 修复:添加真实User-Agent、启用无头模式、随机化等待时间,模拟真实用户行为;检查页面是否出现验证码或拦截提示。

优化后的完整脚本
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait 
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException, NoSuchElementException, ElementClickInterceptedException
import csv
import time

# 浏览器配置
options = webdriver.ChromeOptions()
options.add_experimental_option('excludeSwitches', ['enable-logging'])
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
options.add_argument("--disable-blink-features=AutomationControlled")
options.add_argument("--headless=new")  # 无头模式适合定时任务
options.add_argument("--disable-dev-shm-usage")
options.add_argument("--window-size=1920x1080")

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
wait = WebDriverWait(driver, 15)  # 统一设置等待超时时间

filename = "data.csv"
# 全程保持CSV文件句柄,避免重复IO操作
with open(filename, 'w', encoding='utf-8', newline='') as file:
    writer = csv.writer(file)
    writer.writerow(["URL", "MODEL", "STORAGE", "COLOUR", "GRADE", "PRICE"])

    # 读取链接列表,自动关闭文件
    with open("links.txt", 'r', encoding='utf-8') as link_file:
        line_in_list = [line.strip() for line in link_file if line.strip()]

    for url in line_in_list:
        try:
            # 打开新标签页并切换
            driver.execute_script(f"window.open('{url}','_blank');")
            driver.switch_to.window(driver.window_handles[-1])

            # 处理Cookie弹窗
            try:
                accept_cookie = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Accept all cookies')]")))
                accept_cookie.click()
                time.sleep(1)
            except (TimeoutException, ElementClickInterceptedException):
                print(f"[{url}] Cookie弹窗未找到或无法点击")

            # 获取机型名称
            try:
                model_name = wait.until(EC.visibility_of_element_located((By.XPATH, "//h1[@class='gg-u-text-shout']"))).text
            except TimeoutException:
                print(f"[{url}] 无法获取机型名称")
                model_name = "Unknown"

            # 遍历存储、颜色、等级和价格
            try:
                storage_container = wait.until(EC.visibility_of_all_elements_located((By.XPATH, "//fieldset[@data-testid='memory-selector']/ol/li/label")))
                for storage_item in storage_container:
                    storage_size = storage_item.text.strip()
                    wait.until(EC.element_to_be_clickable(storage_item)).click()
                    time.sleep(1.5)

                    # 遍历颜色选项
                    colour_container = wait.until(EC.visibility_of_all_elements_located((By.XPATH, "//fieldset[@data-testid='colour-selector']/ol/li/label/span")))
                    for colour_item in colour_container:
                        colour = colour_item.text.strip()
                        colour_item.click()
                        time.sleep(1.5)

                        # 遍历等级和价格
                        grade_price_list = wait.until(EC.visibility_of_all_elements_located((By.XPATH, "//div[@class='gg-u-pb-medium']/fieldset/ol/li")))
                        for grade_item in grade_price_list:
                            try:
                                grade = grade_item.find_element(By.XPATH, ".//label/span[1]").text.strip()
                                price = grade_item.find_element(By.XPATH, ".//label/span[2]/span").text.strip()
                                writer.writerow([url, model_name, storage_size, colour, grade, price])
                                print(f"[{url}] 写入数据: {model_name} - {storage_size} - {colour} - {grade} - {price}")
                            except NoSuchElementException:
                                print(f"[{url}] 等级或价格元素缺失")
                                writer.writerow([url, model_name, storage_size, colour, "Unknown", "Unknown"])
            except TimeoutException:
                print(f"[{url}] 机型库存不足或元素加载失败")
                writer.writerow([url, model_name, "None", "None", "Out of Stock", "Out of Stock"])

        except Exception as e:
            print(f"[{url}] 处理失败: {str(e)}")
            writer.writerow([url, "Error", "Error", "Error", "Error", str(e)])
        finally:
            # 关闭当前标签页并切回主窗口
            if len(driver.window_handles) > 1:
                driver.close()
                driver.switch_to.window(driver.window_handles[0])

driver.quit()
print("数据抓取完成")

内容的提问来源于stack exchange,提问作者Reggie18

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 11:35:22