You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium中Cookie弹窗及爬取相关问题求助

Selenium企业信息爬取问题解决方案

问题分析与解决

1. Cookie弹窗延迟等待优化

原代码使用visibility_of_element_located且依赖绝对XPath,容易因页面结构变化或弹窗加载延迟导致等待失效。同时可通过异步线程实现等待弹窗期间启动爬取:

  • 改用相对XPath定位Cookie按钮,避免绝对路径的脆弱性
  • 用element_to_be_clickable替代visibility_of_element_located,确保元素可交互
  • 启动独立线程处理Cookie弹窗,主线程无需等待即可开始爬取

修改后的Cookie处理逻辑:

import threading

def handle_cookies(driver):
    try:
        cookies_button = WebDriverWait(driver, 20).until(
            EC.element_to_be_clickable((By.XPATH, "//div[contains(@class, 'cookie-consent')]//button[contains(text(), 'Accepter')]"))
        )
        cookies_button.click()
    except TimeoutException:
        print("Cookie弹窗未加载,继续爬取")
    except ElementClickInterceptedException:
        driver.execute_script("arguments[0].click();", cookies_button)

# 在scraping函数中启动线程
driver.get(web_adress)
cookie_thread = threading.Thread(target=handle_cookies, args=(driver,))
cookie_thread.start()

2. Excel写入失败修复

原问题核心在于:

  • 绝对XPath定位公司名称失效,find_elements返回空列表
  • 固定从第2行写入,会覆盖已有数据且未处理空行情况
  • find_elements不会抛出NoSuchElementException,原try-except块无效

修改后的爬取写入函数:

def scrape_companies_info(driver, workbook):
    worksheet = workbook.active
    # 改用相对XPath定位公司名称
    company_names = driver.find_elements(By.XPATH, "//div[contains(@class, 'entreprise-item')]/a")
    if not company_names:
        print("未找到公司信息")
        return
    
    # 获取当前最大行,从下一行开始写入
    current_row = worksheet.max_row + 1 if worksheet.max_row >=1 else 2
    for company in company_names:
        company_text = company.text.strip()
        if company_text:
            worksheet.cell(row=current_row, column=1, value=company_text)
            current_row += 1
    workbook.save(excel_spreadsheet)

3. 链接点击滚动优化

原手动滚动30像素的方法不通用,改用以下两种更可靠的方式:

  • JS滚动到可见区域:将元素滚动到视图中心,确保无遮挡
  • ActionChains模拟点击:通过鼠标移动到元素再点击,避免遮挡问题

修改后的链接点击逻辑:

for link in links:
    # 滚动到元素可见位置
    driver.execute_script("arguments[0].scrollIntoView({block: 'center', behavior: 'smooth'});", link)
    # 等待元素可点击
    WebDriverWait(driver, 10).until(EC.element_to_be_clickable(link))
    # 优先用原生点击,失败则用JS点击
    try:
        link.click()
    except ElementClickInterceptedException:
        driver.execute_script("arguments[0].click();", link)
    
    # 等待详情页加载完成
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, "//div[contains(@class, 'entreprise-item')]"))
    )
    scrape_companies_info(driver, workbook)
    driver.back()

最佳实践建议

  • 禁用绝对定位:优先使用元素的class、文本内容、属性等编写相对XPath/CSS选择器,避免页面结构变化导致定位失败
  • 精准等待策略:根据操作选择对应的Expected Conditions,如点击用element_to_be_clickable,元素加载用presence_of_element_located
  • 异步处理非核心操作:Cookie、广告弹窗等不影响核心爬取的操作,用线程异步处理,提升爬取效率
  • Excel写入安全:每次写入前获取当前最大行,避免覆盖已有数据;定期保存工作簿,防止数据丢失
  • 异常细化处理:针对特定异常(如TimeoutException、ElementClickInterceptedException)分别处理,同时记录异常日志便于排查
  • 资源自动清理:使用finally块或with语句确保driver和workbook资源被正确释放

修改后的完整代码

import openpyxl
import threading
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.chrome.options import Options
from selenium.common.exceptions import TimeoutException, ElementClickInterceptedException
from selenium.webdriver.support import expected_conditions as EC

excel_spreadsheet = "companies_scraping.xlsx"
web_adress = "https://infonet.fr/entreprises/liste-des-codes-naf-ape/01-culture-et-production-animale-chasse-et-services-annexes/"

def init_workbook():
    workbook = openpyxl.Workbook()
    worksheet = workbook.active
    headers = ["company_name", "CEO", "APE_code", "department", "creation_date"]
    for col, header in enumerate(headers, 1):
        worksheet.cell(row=1, column=col, value=header)
    workbook.save(excel_spreadsheet)
    return workbook

def init_driver():
    options = Options()
    options.add_argument('--disable-extensions')
    options.add_argument('--disable-popup-blocking')
    options.add_argument('--disable-dev-shm-usage')
    options.add_argument('--no-sandbox')
    options.add_argument('--start-maximized')
    options.add_argument("--incognito")
    chromedriver_path = 'chromedriver' # 替换为你的ChromeDriver路径
    service = Service(chromedriver_path)
    return webdriver.Chrome(service=service, options=options)

def handle_cookies(driver):
    try:
        cookies_button = WebDriverWait(driver, 20).until(
            EC.element_to_be_clickable((By.XPATH, "//div[contains(@class, 'cookie-consent')]//button[contains(text(), 'Accepter')]"))
        )
        cookies_button.click()
    except TimeoutException:
        print("Cookie弹窗未加载")
    except ElementClickInterceptedException:
        driver.execute_script("arguments[0].click();", cookies_button)

def scrape_companies_info(driver, workbook):
    worksheet = workbook.active
    company_names = driver.find_elements(By.XPATH, "//div[contains(@class, 'entreprise-item')]/a")
    if not company_names:
        print("未找到公司信息")
        return
    
    current_row = worksheet.max_row + 1 if worksheet.max_row >=1 else 2
    for company in company_names:
        company_text = company.text.strip()
        if company_text:
            worksheet.cell(row=current_row, column=1, value=company_text)
            current_row += 1
    workbook.save(excel_spreadsheet)
    print(f"写入{len(company_names)}条公司信息")

def scraping(driver, workbook):
    driver.get(web_adress)
    cookie_thread = threading.Thread(target=handle_cookies, args=(driver,))
    cookie_thread.start()

    try:
        links = WebDriverWait(driver, 15).until(
            EC.presence_of_all_elements_located((By.XPATH, "//li[@class='mb-1']/a"))
        )
        print(f"找到{len(links)}个分类链接")
    except TimeoutException:
        print("未找到分类链接")
        links = []

    for index, link in enumerate(links):
        try:
            driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", link)
            WebDriverWait(driver, 10).until(EC.element_to_be_clickable(link))
            link.click()
            
            WebDriverWait(driver, 10).until(
                EC.presence_of_element_located((By.XPATH, "//div[contains(@class, 'entreprise-item')]"))
            )
            scrape_companies_info(driver, workbook)
            driver.back()
            WebDriverWait(driver, 10).until(
                EC.presence_of_all_elements_located((By.XPATH, "//li[@class='mb-1']/a"))
            )
        except ElementClickInterceptedException:
            print(f"第{index+1}个链接点击失败,尝试JS点击")
            driver.execute_script("arguments[0].click();", link)
        except Exception as e:
            print(f"处理第{index+1}个链接出错: {str(e)}")
            driver.back()
            continue

if __name__ == "__main__":
    try:
        workbook = openpyxl.load_workbook(excel_spreadsheet)
        print("加载现有Excel文件")
    except FileNotFoundError:
        workbook = init_workbook()
        print("初始化新Excel文件")

    driver = init_driver()
    try:
        scraping(driver, workbook)
    finally:
        driver.quit()
        workbook.close()
        print("爬取完成,资源已释放")

内容的提问来源于stack exchange,提问作者BobDeTunis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 19:20:00