You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何加速单网站2万+页面的Selenium爬虫抓取?

页面抓取加速方案与代码优化

当前使用Selenium逐页抓取超过20000个页面,现有代码可正常运行但速度极慢,需耗时数日。以下是优化后的代码及核心加速建议,同时保留边抓取边保存数据的逻辑。


原核心循环代码

while True:
    for _ in range(max_retries + 1):
        try:
            section = driver.find_element_by_css_selector('div[id="pageload"]')
            elements = section.find_elements_by_xpath('//li[@class="list-group-item"]/a[contains(., "Lihat")]')
            if not elements:
                break  # No more "Lihat" links, indicating end of data

            for element in elements:
                href = element.get_attribute('href')
                href_set.add(href)

            href_list = list(href_set)
            href_list = [href.replace('/Chome/', '/chome/') for href in href_list]
            with open('output_hrefs.txt', 'w') as file:
                for href in href_list:
                    file.write(href + '\n')

            next_page_link = driver.find_element_by_xpath(f'//ul[@class="pagination pull-left"]/li/a[text()="{page + 1}"]')
            next_page_link.click()
            time.sleep(5)

            page += 1
            break
        except (WebDriverException, NoSuchElementException):
            if _ < max_retries:
                print(f"Retrying page {page}, Attempt {_ + 1}")
                continue
            else:
                print(f"Unable to navigate to page {page} after {max_retries} attempts")
                break

优化后的完整代码

import time
import random
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import WebDriverException, NoSuchElementException, TimeoutException

user_agents = [
    "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/109.0.0.0 Safari/537.36",
    "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_1) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.1 Safari/605.1.15"
]

driver_path = '/users/tosh/downloads/chromedriver'
user_agent = random.choice(user_agents)
chrome_options = Options()

# 启用无头模式,关闭可视化界面
chrome_options.add_argument("--headless=new")
# 禁用图片加载,减少资源消耗
chrome_options.add_argument("--blink-settings=imagesEnabled=false")
# 禁用不必要的浏览器特性
chrome_options.add_argument("--disable-extensions")
chrome_options.add_argument("--disable-gpu")
chrome_options.add_argument("--no-sandbox")
chrome_options.add_argument(f"user-agent={user_agent}")

driver = webdriver.Chrome(executable_path=driver_path, options=chrome_options)
wait = WebDriverWait(driver, 10)  # 显式等待超时时间设为10秒

main_url = 'https://sekolah.data.kemdikbud.go.id/index.php/Chome/pencarian/'
driver.get(main_url)

# 等待搜索按钮可点击后执行点击
wait.until(EC.element_to_be_clickable((By.XPATH, '//*[@id="page-top"]/div[1]/div/div[2]/div[2]/form/div/div/div[3]/div/button'))).click()

max_retries = 3
href_set = set()
page = 1

# 初始化输出文件
with open('output_hrefs.txt', 'w') as file:
    pass

while True:
    success = False
    for attempt in range(max_retries + 1):
        try:
            # 等待目标区域加载完成
            section = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'div[id="pageload"]')))
            # 用相对路径定位元素,避免全局查找浪费资源
            elements = section.find_elements(By.XPATH, './/li[@class="list-group-item"]/a[contains(., "Lihat")]')
            
            if not elements:
                print("No more 'Lihat' links found, stopping crawl.")
                success = True
                break

            # 批量提取新增链接
            new_hrefs = []
            for element in elements:
                href = element.get_attribute('href')
                if href and href not in href_set:
                    href_set.add(href)
                    new_hrefs.append(href.replace('/Chome/', '/chome/'))
            
            # 追加写入新增链接,避免重复写入全部内容
            if new_hrefs:
                with open('output_hrefs.txt', 'a') as file:
                    file.write('\n'.join(new_hrefs) + '\n')
            
            # 等待下一页按钮可点击后跳转
            next_page_link = wait.until(EC.element_to_be_clickable((By.XPATH, f'//ul[@class="pagination pull-left"]/li/a[text()="{page + 1}"]')))
            next_page_link.click()
            
            page += 1
            success = True
            # 随机短等待,规避反爬检测
            time.sleep(random.uniform(1, 3))
            break

        except (WebDriverException, NoSuchElementException, TimeoutException):
            if attempt < max_retries:
                print(f"Retrying page {page}, Attempt {attempt + 1}")
                time.sleep(random.uniform(2, 4))  # 重试前增加随机等待
                continue
            else:
                print(f"Failed to process page {page} after {max_retries} attempts")
                success = False
                break
    if not success or (elements is not None and not elements):
        break

driver.quit()
print("Crawl completed.")

加速抓取的核心建议

1. 替换固定等待为显式等待

原代码中time.sleep(5)是固定等待,无论元素是否加载完成都会浪费时间。改用WebDriverWait仅在元素出现/可点击时继续执行,大幅减少无效等待时长。

2. 启用Chrome无头模式

添加--headless=new参数,浏览器不渲染可视化界面,降低CPU和内存占用,速度提升明显。

3. 禁用不必要的页面资源

禁用图片加载、扩展、GPU加速等非必要资源,减少页面加载的数据量和渲染时间。

4. 优化文件写入逻辑

原代码每次循环重写整个文件,IO开销极大。改为追加写入新增链接,仅写入本次抓取的新内容,减少磁盘IO操作。

5. 并行抓取(进阶方案)

针对20000+页面,单线程效率极低,可使用concurrent.futures.ThreadPoolExecutor实现多线程并行抓取:

  • 每个线程使用独立的浏览器实例
  • 控制并发数(建议5-10线程),避免触发反爬
  • 线程间共享链接集合时需加锁,防止数据冲突

6. 直接调用后端API(最优方案)

打开浏览器开发者工具,分析分页数据的XHR请求,找到后端API接口。用requests库直接请求API,无需渲染页面,速度比Selenium快10-100倍:

  • 解析分页接口的请求参数(如page、limit)
  • 构造请求头模拟浏览器请求
  • 解析JSON响应提取目标链接

7. 优化反爬规避策略

  • 每次请求更换随机User-Agent
  • 使用随机短等待,避免固定间隔触发反爬
  • 实现指数退避重试:重试间隔随失败次数递增(如1s→2s→4s),降低被封禁风险

8. 断点续爬

记录当前抓取的页码到本地文件,下次启动时从上次中断的页码开始,避免重复抓取已完成的页面。

内容的提问来源于stack exchange,提问作者Tosh_gitonga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 12:45:57