如何提升Python中Selenium测试Div组合数据表的代码运行速度?
我完全懂你这种头疼的感觉——用Selenium处理非标准的div数据表本来就麻烦,还遇上频繁远程请求拖慢速度,太影响测试效率了。下面是几个亲测有效的优化方向,帮你减少远程调用、大幅提速脚本:
一次性批量获取元素,避免多次远程请求
每次调用find_element_by_css_selector都会触发一次和浏览器的远程通信,这是速度慢的核心原因之一。换成find_elements_by_css_selector(复数形式)一次性把所有产品元素拉取到本地内存,之后直接在本地遍历处理,就能把N次请求压缩成1次:# 一次性获取所有产品div all_products = driver.find_elements_by_css_selector(".product-container .product-item") # 本地遍历处理每个产品 for product in all_products: name = product.find_element_by_css_selector(".product-name").text price = product.find_element_by_css_selector(".product-price").text # 检查是否有折扣价(用复数形式避免元素不存在时抛异常) discount_price = product.find_elements_by_css_selector(".discount-price") if discount_price: print(f"产品 {name} 折扣价: {discount_price[0].text}")用显式等待替代隐式等待,精准控制等待时长
全局隐式等待会让每一次元素查找都强制等待设定时长(哪怕元素已经存在),白白浪费时间。换成WebDriverWait结合预期条件,只在真正需要等待的场景(比如页面加载、产品列表渲染)等待:from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 等待产品列表加载完成,最多等10秒 all_products = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, ".product-container .product-item")) )这样只有在元素未出现时才会等待,元素加载完成后立即执行后续操作,避免无谓的延迟。
简化CSS选择器,减少DOM遍历开销
复杂的多层嵌套选择器(比如div#main > div.content > div.product-list > div.product-item)会让浏览器花费更多时间遍历DOM,同时增加远程请求的处理耗时。尽量用最精准、简洁的选择器:- 如果产品元素有唯一类名,直接用
.product-item - 带折扣的产品可以用组合类名
.product-item.discount直接定位,不用先找所有产品再判断
简洁的选择器不仅更快,代码也更易读维护。
- 如果产品元素有唯一类名,直接用
直接在浏览器端执行JS批量提取数据
利用driver.execute_script()把数据提取逻辑放到浏览器端执行,只需要一次远程请求就能拿到所有需要的数据,避免多次查找元素+获取属性的重复请求:# 用JS一次性提取所有产品的名称、原价、折扣价 products_data = driver.execute_script(""" return Array.from(document.querySelectorAll('.product-item')).map(item => { const name = item.querySelector('.product-name').textContent.trim(); const originalPrice = item.querySelector('.product-price').textContent.trim(); const discountPriceEl = item.querySelector('.discount-price'); return { name: name, originalPrice: originalPrice, discountPrice: discountPriceEl ? discountPriceEl.textContent.trim() : null }; }) """) # 直接在本地处理返回的数据 for product in products_data: if product['discountPrice']: print(f"{product['name']}: 原价{product['originalPrice']},折扣价{product['discountPrice']}")这种方式的性能提升非常明显,尤其是数据量较大的时候。
启用无头模式,减少浏览器渲染开销
如果测试不需要可视化界面,开启浏览器的无头模式可以大幅减少资源占用,提升运行速度——因为无头模式不需要渲染页面,节省了大量CPU和内存:from selenium.webdriver.chrome.options import Options chrome_options = Options() chrome_options.add_argument("--headless=new") # 新版Chrome无头模式 driver = webdriver.Chrome(options=chrome_options)
这些方法组合起来,应该能显著减少远程请求的数量,让你的脚本运行速度提升一个档次。
内容的提问来源于stack exchange,提问作者Trong Lam Phan

