You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何重构Selenium爬虫函数以提升性能并解决爬取异常?

稳定爬取Rocket League Tracker玩家段位的解决方案

问题根源拆解

  1. 动态加载顺序不稳定:页面异步渲染时,MMR数字元素和段位文本元素的加载顺序不固定,原XPATH定位太宽泛,容易误抓MMR
  2. 资源泄漏:原代码在return后不会执行driver.quit(),导致浏览器进程持续堆积,占用CPU/内存,长期运行会耗尽VM资源导致失效
  3. 超时与重试逻辑缺失:仅依赖单次固定等待,未针对目标元素有效性做循环校验

改进后的实现代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException, NoSuchElementException

def scrape_rank(pName):
    # 配置无头浏览器,降低资源消耗
    options = webdriver.ChromeOptions()
    options.add_argument("--headless=new")
    options.add_argument("--disable-gpu")
    options.add_argument("--no-sandbox")
    options.add_argument("--disable-dev-shm-usage")
    options.add_argument("--disable-extensions")
    
    driver = webdriver.Chrome(options=options)
    wait = WebDriverWait(driver, 15)
    max_attempts = 5
    attempt = 0
    
    try:
        driver.get(f"https://rocketleague.tracker.network/rocket-league/profile/steam/{pName}/overview")
        
        while attempt < max_attempts:
            attempt += 1
            try:
                # 精准定位Standard 3v3的段位元素(可根据页面实际结构微调XPATH)
                rank_element = wait.until(EC.visibility_of_element_located(
                    (By.XPATH, "//div[contains(@class, 'trn-card__content')]//div[text()='Standard 3v3']/following-sibling::div[2]")
                ))
                rank_text = rank_element.text.strip()
                
                if not rank_text.isdigit():
                    return rank_text.split('\n')[0]
                print(f"第{attempt}次抓到MMR,重试中...")
                    
            except (TimeoutException, NoSuchElementException):
                print(f"第{attempt}次定位失败,重试中...")
        
        return "获取段位失败"
        
    finally:
        # 强制释放浏览器资源,避免进程堆积
        driver.quit()

关键优化点说明

  • 资源占用优化:启用无头模式+禁用不必要的浏览器功能,大幅降低Replit和GCP VM上的CPU/内存消耗,解决超时和资源爆满问题
  • 精准元素定位:通过父类trn-card__content缩小范围,指定following-sibling::div[2]直接指向段位文本元素,从根源减少误抓概率
  • 循环校验机制:最多重试5次,每次校验结果是否为纯数字,确保拿到有效段位文本
  • 资源强制释放:finally块保证无论爬取成功或失败,浏览器进程都会被关闭,避免长期运行导致的资源泄漏,解决GCP VM运行一段时间失效的问题

内容的提问来源于stack exchange,提问作者Jared Robertson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 15:43:29