如何加速单网站2万+页面的Selenium爬虫抓取?
页面抓取加速方案与代码优化
当前使用Selenium逐页抓取超过20000个页面,现有代码可正常运行但速度极慢,需耗时数日。以下是优化后的代码及核心加速建议,同时保留边抓取边保存数据的逻辑。
原核心循环代码
while True: for _ in range(max_retries + 1): try: section = driver.find_element_by_css_selector('div[id="pageload"]') elements = section.find_elements_by_xpath('//li[@class="list-group-item"]/a[contains(., "Lihat")]') if not elements: break # No more "Lihat" links, indicating end of data for element in elements: href = element.get_attribute('href') href_set.add(href) href_list = list(href_set) href_list = [href.replace('/Chome/', '/chome/') for href in href_list] with open('output_hrefs.txt', 'w') as file: for href in href_list: file.write(href + '\n') next_page_link = driver.find_element_by_xpath(f'//ul[@class="pagination pull-left"]/li/a[text()="{page + 1}"]') next_page_link.click() time.sleep(5) page += 1 break except (WebDriverException, NoSuchElementException): if _ < max_retries: print(f"Retrying page {page}, Attempt {_ + 1}") continue else: print(f"Unable to navigate to page {page} after {max_retries} attempts") break
优化后的完整代码
import time import random from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import WebDriverException, NoSuchElementException, TimeoutException user_agents = [ "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/109.0.0.0 Safari/537.36", "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_1) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.1 Safari/605.1.15" ] driver_path = '/users/tosh/downloads/chromedriver' user_agent = random.choice(user_agents) chrome_options = Options() # 启用无头模式,关闭可视化界面 chrome_options.add_argument("--headless=new") # 禁用图片加载,减少资源消耗 chrome_options.add_argument("--blink-settings=imagesEnabled=false") # 禁用不必要的浏览器特性 chrome_options.add_argument("--disable-extensions") chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument(f"user-agent={user_agent}") driver = webdriver.Chrome(executable_path=driver_path, options=chrome_options) wait = WebDriverWait(driver, 10) # 显式等待超时时间设为10秒 main_url = 'https://sekolah.data.kemdikbud.go.id/index.php/Chome/pencarian/' driver.get(main_url) # 等待搜索按钮可点击后执行点击 wait.until(EC.element_to_be_clickable((By.XPATH, '//*[@id="page-top"]/div[1]/div/div[2]/div[2]/form/div/div/div[3]/div/button'))).click() max_retries = 3 href_set = set() page = 1 # 初始化输出文件 with open('output_hrefs.txt', 'w') as file: pass while True: success = False for attempt in range(max_retries + 1): try: # 等待目标区域加载完成 section = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'div[id="pageload"]'))) # 用相对路径定位元素,避免全局查找浪费资源 elements = section.find_elements(By.XPATH, './/li[@class="list-group-item"]/a[contains(., "Lihat")]') if not elements: print("No more 'Lihat' links found, stopping crawl.") success = True break # 批量提取新增链接 new_hrefs = [] for element in elements: href = element.get_attribute('href') if href and href not in href_set: href_set.add(href) new_hrefs.append(href.replace('/Chome/', '/chome/')) # 追加写入新增链接,避免重复写入全部内容 if new_hrefs: with open('output_hrefs.txt', 'a') as file: file.write('\n'.join(new_hrefs) + '\n') # 等待下一页按钮可点击后跳转 next_page_link = wait.until(EC.element_to_be_clickable((By.XPATH, f'//ul[@class="pagination pull-left"]/li/a[text()="{page + 1}"]'))) next_page_link.click() page += 1 success = True # 随机短等待,规避反爬检测 time.sleep(random.uniform(1, 3)) break except (WebDriverException, NoSuchElementException, TimeoutException): if attempt < max_retries: print(f"Retrying page {page}, Attempt {attempt + 1}") time.sleep(random.uniform(2, 4)) # 重试前增加随机等待 continue else: print(f"Failed to process page {page} after {max_retries} attempts") success = False break if not success or (elements is not None and not elements): break driver.quit() print("Crawl completed.")
加速抓取的核心建议
1. 替换固定等待为显式等待
原代码中time.sleep(5)是固定等待,无论元素是否加载完成都会浪费时间。改用WebDriverWait仅在元素出现/可点击时继续执行,大幅减少无效等待时长。
2. 启用Chrome无头模式
添加--headless=new参数,浏览器不渲染可视化界面,降低CPU和内存占用,速度提升明显。
3. 禁用不必要的页面资源
禁用图片加载、扩展、GPU加速等非必要资源,减少页面加载的数据量和渲染时间。
4. 优化文件写入逻辑
原代码每次循环重写整个文件,IO开销极大。改为追加写入新增链接,仅写入本次抓取的新内容,减少磁盘IO操作。
5. 并行抓取(进阶方案)
针对20000+页面,单线程效率极低,可使用concurrent.futures.ThreadPoolExecutor实现多线程并行抓取:
- 每个线程使用独立的浏览器实例
- 控制并发数(建议5-10线程),避免触发反爬
- 线程间共享链接集合时需加锁,防止数据冲突
6. 直接调用后端API(最优方案)
打开浏览器开发者工具,分析分页数据的XHR请求,找到后端API接口。用requests库直接请求API,无需渲染页面,速度比Selenium快10-100倍:
- 解析分页接口的请求参数(如
page、limit) - 构造请求头模拟浏览器请求
- 解析JSON响应提取目标链接
7. 优化反爬规避策略
- 每次请求更换随机User-Agent
- 使用随机短等待,避免固定间隔触发反爬
- 实现指数退避重试:重试间隔随失败次数递增(如1s→2s→4s),降低被封禁风险
8. 断点续爬
记录当前抓取的页码到本地文件,下次启动时从上次中断的页码开始,避免重复抓取已完成的页面。
内容的提问来源于stack exchange,提问作者Tosh_gitonga
相关产品推荐
相关产品推荐

