You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium+Python循环迭代中表格抓取变慢的问题排查咨询

问题:Selenium表格抓取耗时随迭代递增的排查与优化

问题背景

运行以下Selenium表格抓取代码时,功能正常但仅表格抓取环节耗时随迭代次数逐渐增加,初始不到10秒,后续会增至3分钟以上,循环其他部分执行时间稳定。

核心抓取代码

while (True):
   try:
        time.sleep(5)
        start_time = time.time() 
        # Get the scrapped table data
        table_data = []
        rows = table.find_elements(By.TAG_NAME, 'tr')
        for row in rows:
            row_data = []
            cells = row.find_elements(By.TAG_NAME, 'td')
            for cell in cells:
                row_data.append(cell.text)
            table_data.append(row_data)
            print(row_data)
        print("--- %s seconds ---" % (time.time() - start_time)) 
###
# then saves to the database
# changes to a page with the next table and    
# return to while loop start
###

耗时数据(单位:秒)

--- 9.43001103401184 seconds ---
--- 9.989665746688843 seconds ---
--- 10.980393886566162 seconds ---
--- 12.015849828720093 seconds ---
--- 13.079446077346802 seconds ---
--- 14.133357286453247 seconds ---
--- 15.820979833602905 seconds ---
--- 16.55774736404419 seconds ---
--- 19.471170663833618 seconds ---
--- 25.035650730133057 seconds ---
--- 27.544485092163086 seconds ---
--- 30.97364568710327 seconds ---
--- 30.9657781124115 seconds ---
--- 34.56339645385742 seconds ---
--- 35.829299211502075 seconds ---

已排查情况

  • 可一次性获取完整HTML无延迟,排除网络或页面加载问题
  • 任务管理器显示内存占用正常,但CPU活跃度较高
  • 代码可连续运行20小时以上不崩溃

Selenium初始化代码

from selenium import webdriver
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.common.by import By

nav = webdriver.Chrome()
nav.get("https://uri_from_where_i_get_test_data.com/page")

原因分析

  1. 元素引用累积:每次循环中find_elements获取的元素对象会保留对浏览器DOM的引用,随着迭代次数增加,未清理的引用会导致浏览器内部DOM缓存膨胀,查询效率下降。
  2. 逐元素定位损耗:嵌套循环中每次调用find_elements(By.TAG_NAME, 'td')都会发起新的DOM查询,多次重复查询会累积耗时,尤其表格行数/列数较多时。
  3. ChromeDriver隐性泄漏:长期运行下,ChromeDriver可能存在隐性内存泄漏,虽系统内存显示正常,但浏览器进程内部资源占用会逐渐增加,影响DOM操作速度。

优化解决方案

1. 一次性提取HTML用解析库处理(推荐)

跳过Selenium逐元素查询,直接获取表格HTML源码,用BeautifulSoup解析,大幅减少DOM交互次数:

from bs4 import BeautifulSoup
import time

while (True):
    try:
        time.sleep(5)
        start_time = time.time()
        # 直接获取表格完整HTML
        table_html = table.get_attribute('outerHTML')
        soup = BeautifulSoup(table_html, 'lxml')
        table_data = []
        # 批量解析表格数据
        for row in soup.find_all('tr'):
            row_data = [cell.get_text(strip=True) for cell in row.find_all('td')]
            table_data.append(row_data)
            print(row_data)
        print("--- %s seconds ---" % (time.time() - start_time))
        # 后续数据库存储、页面切换逻辑...
    except Exception as e:
        print(e)

2. 显式清理元素引用

每次循环结束后删除元素列表并触发垃圾回收,减少浏览器DOM引用累积:

import gc
import time

while (True):
    try:
        time.sleep(5)
        start_time = time.time()
        table_data = []
        rows = table.find_elements(By.TAG_NAME, 'tr')
        for row in rows:
            row_data = []
            cells = row.find_elements(By.TAG_NAME, 'td')
            for cell in cells:
                row_data.append(cell.text)
            table_data.append(row_data)
            print(row_data)
            # 清理当前行的单元格引用
            del cells
        print("--- %s seconds ---" % (time.time() - start_time))
        # 清理行引用并触发垃圾回收
        del rows
        gc.collect()
        # 后续数据库存储、页面切换逻辑...
    except Exception as e:
        print(e)

3. 定期重启浏览器进程

针对ChromeDriver隐性泄漏,每N次循环后重启浏览器重置环境:

import time

loop_count = 0
MAX_LOOPS_BEFORE_RESTART = 50  # 根据实际场景调整阈值

while (True):
    try:
        loop_count +=1
        time.sleep(5)
        start_time = time.time()
        # 表格抓取逻辑(原代码)
        table_data = []
        rows = table.find_elements(By.TAG_NAME, 'tr')
        for row in rows:
            row_data = []
            cells = row.find_elements(By.TAG_NAME, 'td')
            for cell in cells:
                row_data.append(cell.text)
            table_data.append(row_data)
            print(row_data)
        print("--- %s seconds ---" % (time.time() - start_time))
        
        # 达到阈值时重启浏览器
        if loop_count >= MAX_LOOPS_BEFORE_RESTART:
            nav.quit()
            # 重新初始化浏览器
            nav = webdriver.Chrome()
            nav.get("https://uri_from_where_i_get_test_data.com/page")
            loop_count = 0
            
        # 后续数据库存储、页面切换逻辑...
    except Exception as e:
        print(e)

4. 优化Selenium启动参数

添加Chrome启动参数,禁用不必要功能减少资源占用:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

chrome_options = Options()
# 禁用图片加载
chrome_options.add_argument('--blink-settings=imagesEnabled=false')
# 禁用GPU加速
chrome_options.add_argument('--disable-gpu')
# 启用无头模式(可选,减少UI资源消耗)
# chrome_options.add_argument('--headless=new')
# 禁用JavaScript(仅当表格无需JS渲染时使用)
# chrome_options.add_argument('--disable-javascript')

nav = webdriver.Chrome(options=chrome_options)
nav.get("https://uri_from_where_i_get_test_data.com/page")

内容的提问来源于stack exchange,提问作者Henrique Vilela

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 06:37:30