Python Selenium爬取Diretta页面过慢,wbe.get耗时久如何解决?
解决Selenium爬取Diretta页面
get()方法耗时过长的问题 问题背景
需要爬取https://www.diretta.it/partita/ALdJKzeJ/#/classifiche/table/overall页面数据,使用Selenium实现时单次执行耗时达30-41秒,通过时间打点测试发现**wbe.get(url)环节耗时最久**(从13:28:32到13:29:01,接近29秒),多次执行后累积耗时过高。
用户的打点测试代码及输出如下:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from time import perf_counter import datetime chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--headless") service = Service('/Users/paolopiceni/Desktop/scraping/chromedriver') XPATH = '//*[@class="ui-table__row table__row--selected "]/div[@class="table__cell table__cell--participant "]' HC = 'duelParticipant__home' AC = 'duelParticipant__away' IW = 5 def main(): with webdriver.Chrome('/Users/paolopiceni/Desktop/scraping/chromedriver', options=chrome_options) as wbe: wbe.implicitly_wait(IW) print(datetime.datetime.now()) wbe.get("https://www.diretta.it/partita/ALdJKzeJ/#/classifiche/table/overall") print(datetime.datetime.now()) TEAM_HOME = wbe.find_element(By.CLASS_NAME, HC).text TEAM_AWAY = wbe.find_element(By.CLASS_NAME, AC).text print(datetime.datetime.now()) teams_class_totale = wbe.find_elements(By.XPATH, XPATH) print(datetime.datetime.now()) assert len(teams_class_totale) > 1 return TEAM_HOME, TEAM_AWAY, teams_class_totale[0].text, teams_class_totale[1].text if __name__ == '__main__': start = perf_counter() print(*main()) end = perf_counter() print(f'Duration={end-start:.4}s')
运行输出:
2023-06-30 13:28:32.871837 2023-06-30 13:29:01.530982 2023-06-30 13:29:04.029629 2023-06-30 13:29:10.085035 El Gaish Ceramica Cleopatra Ceramica Cleopatra El Gaish Duration=41.54s
耗时原因分析
- 页面冗余资源加载:Diretta作为体育数据网站,会加载大量非必要资源(广告脚本、视频组件、实时推送服务、第三方统计代码等),Selenium默认等待所有资源加载完成才会执行后续逻辑,这些资源的加载直接拉长了
get()方法的耗时。 - 反爬机制限制:网站可能检测到无头浏览器请求,故意延长页面加载时间或加入隐性验证环节,导致加载缓慢。
- 无头模式性能配置不足:默认的无头Chrome没有启用性能优化参数,比如未禁用图片、CSS加载,进一步增加了资源加载时间。
- 默认页面加载策略:Selenium默认使用
normal加载策略,需等待所有资源(包括静态资源、异步请求)加载完成才会结束get()方法。
优化解决方案
1. 调整页面加载策略
将加载策略改为eager,即DOM结构加载完成后就停止等待,无需等待所有静态资源加载:
chrome_options.page_load_strategy = 'eager'
2. 禁用非必要资源加载
通过Chrome配置禁止加载图片、CSS、插件等非必要资源,减少页面加载量:
prefs = { "profile.managed_default_content_settings.images": 2, "profile.managed_default_content_settings.stylesheets": 2, "profile.managed_default_content_settings.plugins": 2, "profile.managed_default_content_settings.popups": 2 } chrome_options.add_experimental_option("prefs", prefs)
3. 优化无头模式配置
使用新版无头模式并添加性能优化参数:
chrome_options.add_argument("--headless=new") # Chrome 112+支持的新版无头模式,性能更优 chrome_options.add_argument("--disable-dev-shm-usage") # 解决内存限制问题 chrome_options.add_argument("--disable-extensions") chrome_options.add_argument("--disable-background-networking")
4. 用显式等待替代隐式等待
避免全局隐式等待带来的不必要延迟,仅等待目标元素加载完成:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC wait = WebDriverWait(wbe, 5) TEAM_HOME = wait.until(EC.presence_of_element_located((By.CLASS_NAME, HC))).text TEAM_AWAY = wait.until(EC.presence_of_element_located((By.CLASS_NAME, AC))).text teams_class_totale = wait.until(EC.presence_of_all_elements_located((By.XPATH, XPATH)))
5. 尝试静态爬取(如果可行)
若页面目标数据是静态渲染的,直接用requests+BeautifulSoup替代Selenium,大幅降低耗时:
import requests from bs4 import BeautifulSoup headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } url = "https://www.diretta.it/partita/ALdJKzeJ/#/classifiche/table/overall" response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') team_home = soup.find(class_='duelParticipant__home').text.strip() team_away = soup.find(class_='duelParticipant__away').text.strip()
优化后的完整示例代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from time import perf_counter def main(): chrome_options = webdriver.ChromeOptions() # 设置页面加载策略为eager chrome_options.page_load_strategy = 'eager' # 新版无头模式及性能优化参数 chrome_options.add_argument("--headless=new") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") # 禁用非必要资源 prefs = { "profile.managed_default_content_settings.images": 2, "profile.managed_default_content_settings.stylesheets": 2, "profile.managed_default_content_settings.plugins": 2 } chrome_options.add_experimental_option("prefs", prefs) with webdriver.Chrome(options=chrome_options) as wbe: wait = WebDriverWait(wbe, 5) wbe.get("https://www.diretta.it/partita/ALdJKzeJ/#/classifiche/table/overall") TEAM_HOME = wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'duelParticipant__home'))).text TEAM_AWAY = wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'duelParticipant__away'))).text teams_class_totale = wait.until(EC.presence_of_all_elements_located((By.XPATH, '//*[@class="ui-table__row table__row--selected "]/div[@class="table__cell table__cell--participant "]'))) assert len(teams_class_totale) > 1 return TEAM_HOME, TEAM_AWAY, teams_class_totale[0].text, teams_class_totale[1].text if __name__ == '__main__': start = perf_counter() print(*main()) end = perf_counter() print(f'Duration={end-start:.4}s')
内容的提问来源于stack exchange,提问作者mario
相关产品推荐
相关产品推荐

