如何重构Selenium爬虫函数以提升性能并解决爬取异常?
稳定爬取Rocket League Tracker玩家段位的解决方案
问题根源拆解
- 动态加载顺序不稳定:页面异步渲染时,MMR数字元素和段位文本元素的加载顺序不固定,原XPATH定位太宽泛,容易误抓MMR
- 资源泄漏:原代码在
return后不会执行driver.quit(),导致浏览器进程持续堆积,占用CPU/内存,长期运行会耗尽VM资源导致失效 - 超时与重试逻辑缺失:仅依赖单次固定等待,未针对目标元素有效性做循环校验
改进后的实现代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException, NoSuchElementException def scrape_rank(pName): # 配置无头浏览器,降低资源消耗 options = webdriver.ChromeOptions() options.add_argument("--headless=new") options.add_argument("--disable-gpu") options.add_argument("--no-sandbox") options.add_argument("--disable-dev-shm-usage") options.add_argument("--disable-extensions") driver = webdriver.Chrome(options=options) wait = WebDriverWait(driver, 15) max_attempts = 5 attempt = 0 try: driver.get(f"https://rocketleague.tracker.network/rocket-league/profile/steam/{pName}/overview") while attempt < max_attempts: attempt += 1 try: # 精准定位Standard 3v3的段位元素(可根据页面实际结构微调XPATH) rank_element = wait.until(EC.visibility_of_element_located( (By.XPATH, "//div[contains(@class, 'trn-card__content')]//div[text()='Standard 3v3']/following-sibling::div[2]") )) rank_text = rank_element.text.strip() if not rank_text.isdigit(): return rank_text.split('\n')[0] print(f"第{attempt}次抓到MMR,重试中...") except (TimeoutException, NoSuchElementException): print(f"第{attempt}次定位失败,重试中...") return "获取段位失败" finally: # 强制释放浏览器资源,避免进程堆积 driver.quit()
关键优化点说明
- 资源占用优化:启用无头模式+禁用不必要的浏览器功能,大幅降低Replit和GCP VM上的CPU/内存消耗,解决超时和资源爆满问题
- 精准元素定位:通过父类
trn-card__content缩小范围,指定following-sibling::div[2]直接指向段位文本元素,从根源减少误抓概率 - 循环校验机制:最多重试5次,每次校验结果是否为纯数字,确保拿到有效段位文本
- 资源强制释放:
finally块保证无论爬取成功或失败,浏览器进程都会被关闭,避免长期运行导致的资源泄漏,解决GCP VM运行一段时间失效的问题
内容的提问来源于stack exchange,提问作者Jared Robertson
相关产品推荐
相关产品推荐

