You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Selenium爬虫,提速采集nafdac分页表格数据

Selenium分页采集效率优化方案

原有代码核心问题

  • 逻辑错误:将下一页点击逻辑写在单条数据遍历循环内,每采集1行就触发1次下页跳转+10秒硬等待,还额外在每行采集前加了10秒等待,仅这一项就导致采集效率比正常逻辑慢了几十倍
  • 冗余硬等待:大量使用time.sleep()固定等待,即使页面提前加载完成也必须等满设定时间,浪费大量时间
  • 浏览器无优化:默认启动带UI的Chrome,会加载图片、冗余样式、无用插件等非必要资源,拖慢页面渲染速度
  • 交互频率过高:逐行逐单元格调用Selenium接口读取内容,和浏览器的交互次数过多,额外增加耗时

具体优化措施

基础优化(基于Selenium框架,改造成本极低)

  1. 修复逻辑错误:把下一页点击逻辑移动到每页所有数据采集完成之后,每页仅点击1次下页
  2. 移除所有硬等待:全部替换为条件等待,仅等待表格/下页按钮渲染完成就执行下一步操作,不需要固定等待
  3. 配置Chrome无头模式+性能参数:禁用UI、图片加载、GPU加速、插件等非必要功能,大幅降低浏览器资源消耗
  4. 批量读取表格数据:一次性提取当前页所有单元格内容,减少Selenium和浏览器的交互次数

进阶优化(效率提升10倍以上)

直接抓取分页AJAX接口:该网站分页是动态请求,切换分页时会向后端发送带页码参数的请求,直接模拟请求接口获取JSON数据,不需要启动浏览器,是最高效的采集方案。


优化后Selenium代码示例

# pip install webdriver-manager pandas selenium --user
from selenium import webdriver
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.common.exceptions import NoSuchElementException, StaleElementReferenceException
from selenium.webdriver.support import expected_conditions as ec
import pandas as pd

# 配置Chrome性能优化参数
chrome_options = webdriver.ChromeOptions()
# 无头模式(无UI)
chrome_options.add_argument("--headless=new")
# 禁用图片加载
chrome_options.add_argument("--blink-settings=imagesEnabled=false")
# 禁用GPU加速
chrome_options.add_argument("--disable-gpu")
# 禁用插件
chrome_options.add_argument("--disable-plugins")
# 关闭沙箱模式
chrome_options.add_argument("--no-sandbox")
# 内存优化
chrome_options.add_argument("--disable-dev-shm-usage")

driver = webdriver.Chrome(ChromeDriverManager().install(), options=chrome_options)
driver.get('https://www.nafdac.gov.ng/our-services/registered-products/')

container2 = []
wait_time_out = 10
ignored_exceptions = (NoSuchElementException, StaleElementReferenceException,)
total_page = 5802

for page in range(total_page):
    # 等待当前页表格加载完成
    rows = WebDriverWait(driver, wait_time_out, ignored_exceptions=ignored_exceptions).until(
        ec.presence_of_all_elements_located((By.CSS_SELECTOR, '#table_1 tbody tr'))
    )
    # 批量采集当前页所有数据
    for row in rows:
        container2.append([td.text for td in row.find_elements(By.TAG_NAME, 'td')])
    # 非最后一页时点击下一页
    if page < total_page - 1:
        # 等待下一页按钮可点击后跳转
        next_btn = WebDriverWait(driver, wait_time_out, ignored_exceptions=ignored_exceptions).until(
            ec.element_to_be_clickable((By.CSS_SELECTOR, '#table_1_next'))
        )
        next_btn.click()
        # 等待表格更新完成(判断第一行内容变化,避免采集重复数据)
        WebDriverWait(driver, wait_time_out, ignored_exceptions=ignored_exceptions).until(
            ec.staleness_of(rows[0])
        )

# 导出数据
pd.DataFrame(container2).to_csv('registered_products.csv', index=False, encoding='utf-8-sig')
driver.quit()

内容的提问来源于stack exchange,提问作者Uyiosa Akpasubi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 04:06:05