Selenium+Python循环迭代中表格抓取变慢的问题排查咨询
问题:Selenium表格抓取耗时随迭代递增的排查与优化
问题背景
运行以下Selenium表格抓取代码时,功能正常但仅表格抓取环节耗时随迭代次数逐渐增加,初始不到10秒,后续会增至3分钟以上,循环其他部分执行时间稳定。
核心抓取代码
while (True): try: time.sleep(5) start_time = time.time() # Get the scrapped table data table_data = [] rows = table.find_elements(By.TAG_NAME, 'tr') for row in rows: row_data = [] cells = row.find_elements(By.TAG_NAME, 'td') for cell in cells: row_data.append(cell.text) table_data.append(row_data) print(row_data) print("--- %s seconds ---" % (time.time() - start_time)) ### # then saves to the database # changes to a page with the next table and # return to while loop start ###
耗时数据(单位:秒)
--- 9.43001103401184 seconds --- --- 9.989665746688843 seconds --- --- 10.980393886566162 seconds --- --- 12.015849828720093 seconds --- --- 13.079446077346802 seconds --- --- 14.133357286453247 seconds --- --- 15.820979833602905 seconds --- --- 16.55774736404419 seconds --- --- 19.471170663833618 seconds --- --- 25.035650730133057 seconds --- --- 27.544485092163086 seconds --- --- 30.97364568710327 seconds --- --- 30.9657781124115 seconds --- --- 34.56339645385742 seconds --- --- 35.829299211502075 seconds ---
已排查情况
- 可一次性获取完整HTML无延迟,排除网络或页面加载问题
- 任务管理器显示内存占用正常,但CPU活跃度较高
- 代码可连续运行20小时以上不崩溃
Selenium初始化代码
from selenium import webdriver from selenium.webdriver.common.keys import Keys from selenium.webdriver.common.by import By nav = webdriver.Chrome() nav.get("https://uri_from_where_i_get_test_data.com/page")
原因分析
- 元素引用累积:每次循环中
find_elements获取的元素对象会保留对浏览器DOM的引用,随着迭代次数增加,未清理的引用会导致浏览器内部DOM缓存膨胀,查询效率下降。 - 逐元素定位损耗:嵌套循环中每次调用
find_elements(By.TAG_NAME, 'td')都会发起新的DOM查询,多次重复查询会累积耗时,尤其表格行数/列数较多时。 - ChromeDriver隐性泄漏:长期运行下,ChromeDriver可能存在隐性内存泄漏,虽系统内存显示正常,但浏览器进程内部资源占用会逐渐增加,影响DOM操作速度。
优化解决方案
1. 一次性提取HTML用解析库处理(推荐)
跳过Selenium逐元素查询,直接获取表格HTML源码,用BeautifulSoup解析,大幅减少DOM交互次数:
from bs4 import BeautifulSoup import time while (True): try: time.sleep(5) start_time = time.time() # 直接获取表格完整HTML table_html = table.get_attribute('outerHTML') soup = BeautifulSoup(table_html, 'lxml') table_data = [] # 批量解析表格数据 for row in soup.find_all('tr'): row_data = [cell.get_text(strip=True) for cell in row.find_all('td')] table_data.append(row_data) print(row_data) print("--- %s seconds ---" % (time.time() - start_time)) # 后续数据库存储、页面切换逻辑... except Exception as e: print(e)
2. 显式清理元素引用
每次循环结束后删除元素列表并触发垃圾回收,减少浏览器DOM引用累积:
import gc import time while (True): try: time.sleep(5) start_time = time.time() table_data = [] rows = table.find_elements(By.TAG_NAME, 'tr') for row in rows: row_data = [] cells = row.find_elements(By.TAG_NAME, 'td') for cell in cells: row_data.append(cell.text) table_data.append(row_data) print(row_data) # 清理当前行的单元格引用 del cells print("--- %s seconds ---" % (time.time() - start_time)) # 清理行引用并触发垃圾回收 del rows gc.collect() # 后续数据库存储、页面切换逻辑... except Exception as e: print(e)
3. 定期重启浏览器进程
针对ChromeDriver隐性泄漏,每N次循环后重启浏览器重置环境:
import time loop_count = 0 MAX_LOOPS_BEFORE_RESTART = 50 # 根据实际场景调整阈值 while (True): try: loop_count +=1 time.sleep(5) start_time = time.time() # 表格抓取逻辑(原代码) table_data = [] rows = table.find_elements(By.TAG_NAME, 'tr') for row in rows: row_data = [] cells = row.find_elements(By.TAG_NAME, 'td') for cell in cells: row_data.append(cell.text) table_data.append(row_data) print(row_data) print("--- %s seconds ---" % (time.time() - start_time)) # 达到阈值时重启浏览器 if loop_count >= MAX_LOOPS_BEFORE_RESTART: nav.quit() # 重新初始化浏览器 nav = webdriver.Chrome() nav.get("https://uri_from_where_i_get_test_data.com/page") loop_count = 0 # 后续数据库存储、页面切换逻辑... except Exception as e: print(e)
4. 优化Selenium启动参数
添加Chrome启动参数,禁用不必要功能减少资源占用:
from selenium import webdriver from selenium.webdriver.chrome.options import Options chrome_options = Options() # 禁用图片加载 chrome_options.add_argument('--blink-settings=imagesEnabled=false') # 禁用GPU加速 chrome_options.add_argument('--disable-gpu') # 启用无头模式(可选,减少UI资源消耗) # chrome_options.add_argument('--headless=new') # 禁用JavaScript(仅当表格无需JS渲染时使用) # chrome_options.add_argument('--disable-javascript') nav = webdriver.Chrome(options=chrome_options) nav.get("https://uri_from_where_i_get_test_data.com/page")
内容的提问来源于stack exchange,提问作者Henrique Vilela
相关产品推荐
相关产品推荐

