You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化nafdac.gov.ng站点表格数据Selenium爬虫的爬取速度

爬取效率优化方案

问题背景

目标爬取站点为 nafdac.gov.ng/our-services/registered-products,站点整体结构固定,仅每页表格内容随翻页操作更新。原Selenium代码爬取200页耗时7小时,全站共5802页,需要优化爬取效率。

原始代码问题诊断

你提供的原始代码如下:

# pip install webdriver-manager --user
from selenium import webdriver
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.common.exceptions import NoSuchElementException, StaleElementReferenceException
from selenium.webdriver.support import expected_conditions as ec
import pandas as pd
import time

driver = webdriver.Chrome(ChromeDriverManager().install())
driver.get('https://www.nafdac.gov.ng/our-services/registered-products/')

container2 = []
wait_time_out = 20
ignored_exceptions = (NoSuchElementException,StaleElementReferenceException,)

for _ in range(0, 5802+1):
    rows = WebDriverWait(driver, wait_time_out, ignored_exceptions=ignored_exceptions).until(
        ec.presence_of_all_elements_located((By.XPATH, '//*[@id="table_1"]/tbody/tr'))
    )
    for row in rows:
        time.sleep(10) 
    container2.append([table_data.text for table_data in row.find_elements(By.TAG_NAME, 'td')])
    WebDriverWait(driver, wait_time_out, ignored_exceptions=ignored_exceptions).until(
        ec.presence_of_element_located((By.XPATH, '//*[@id="table_1_next"]'))
    ).click()
    time.sleep(10) 

原代码最大的耗时来源是两处无意义的硬等待,每行数据遍历要等10秒,翻页后又等10秒,仅等待时间就占了总耗时的99%以上。

具体优化方案

  • 移除所有不必要的time.sleep()硬等待,所有等待逻辑用Selenium自带的显式等待实现,仅在必要的反爬场景下加1到2秒的短等待即可。
  • 开启Chrome无头模式,关闭GUI渲染,大幅降低资源消耗,提升加载速度,启动driver时新增如下配置:
chrome_options = webdriver.ChromeOptions()
# 开启无头模式
chrome_options.add_argument('--headless=new')
# 禁用GPU加速
chrome_options.add_argument('--disable-gpu')
# 禁用沙箱模式
chrome_options.add_argument('--no-sandbox')
# 禁用图片加载
chrome_options.add_argument('--blink-settings=imagesEnabled=false')
# 禁用CSS加载
chrome_options.add_experimental_option("prefs", {
    "profile.managed_default_content_settings.images": 2,
    "profile.default_content_setting_values.css": 2
})
driver = webdriver.Chrome(ChromeDriverManager().install(), options=chrome_options)
  • 用JS批量提取表格数据,避免逐行逐元素调用Selenium接口,减少跨进程通信开销:
# 直接用JS一次性提取整个表格内容,比逐行遍历快10倍以上
page_data = driver.execute_script("""
    const rows = document.querySelectorAll('#table_1 tbody tr');
    return Array.from(rows).map(row => Array.from(row.querySelectorAll('td')).map(cell => cell.textContent.trim()));
""")
container2.extend(page_data)
  • 翻页后新增表格内容更新判断,避免爬取重复数据:翻页后等待新页面的第一行表格内容和上一页第一行内容不同,再开始提取数据,比固定等待10秒高效很多。
  • 优先尝试抓接口请求:打开浏览器F12开发者工具的网络面板,点击翻页查看XHR类型的请求,找到表格数据对应的后台接口,直接用requests库请求接口获取JSON数据,不需要渲染页面,效率比Selenium至少高20倍,合理控制请求间隔的情况下,5802页仅需要1到2小时就能爬取完成。
  • 新增异常重试逻辑,单页加载失败最多重试3次,避免长时间等待超时浪费时间。

内容的提问来源于stack exchange,提问作者Uyiosa Akpasubi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 16:45:02