You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Colab中Selenium访问部分URL超时,顺序影响执行结果求助

解决CMS爬虫在Google Colab中的TimeoutException问题

问题分析

你的爬虫在本地运行正常,但在Colab中连续访问第二个URL时触发driver.get(url)超时,单独访问该URL却能成功。这大概率是因为Colab共享环境下,同一个Chrome会话的缓存、Cookie或连续请求触发了CMS网站的反爬机制,导致后续请求被拦截。

解决方案

1. 为每个URL创建独立的浏览器实例

复用同一个浏览器会话可能会携带之前的会话信息,被目标网站识别为异常请求。修改代码,在循环内初始化新的Chrome实例,每个URL使用独立会话:

from selenium import webdriver
from selenium.common.exceptions import TimeoutException
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.wait import WebDriverWait


def download_documents() -> None:
    """Download billing code documents from CMS"""

    chrome_options = Options()
    chrome_options.add_argument("--headless")
    chrome_options.add_argument("--no-sandbox")
    chrome_options.add_argument("--disable-dev-shm-usage")
    # 添加反自动化检测参数
    chrome_options.add_argument("--disable-blink-features=AutomationControlled")
    chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
    chrome_options.add_experimental_option('useAutomationExtension', False)

    working_url = "https://www.cms.gov/medicare-coverage-database/view/article.aspx?articleid=59626&ver=6"
    not_working_url = "https://www.cms.gov/medicare-coverage-database/view/lcd.aspx?lcdid=36377&ver=19"

    for row in [working_url, not_working_url]:
        print(f"Retrieving from {row}...")
        # 每个URL新建浏览器实例
        driver = webdriver.Chrome(options=chrome_options)
        try:
            driver.get(row)

            print("Wait for webdriver...")
            wait = WebDriverWait(driver, 2)

            print("Attempting license accept...")
            try:
                wait.until(EC.element_to_be_clickable((By.ID, "btnAcceptLicense"))).click()
            except TimeoutException:
                pass
            
            wait = WebDriverWait(driver, 4)
            print("Attempting pop up close...")
            try:
                wait.until(
                    EC.element_to_be_clickable(
                        (By.XPATH, "//button[@data-page-action='Clicked the Tracking Sheet Close button.']")
                    )
                ).click()
            except TimeoutException:
                pass
            
            print("Attempting download...")
            driver.find_element(By.ID, "btnDownload").click()
        finally:
            # 确保每个实例使用后关闭
            driver.quit()

download_documents()

2. 添加请求间隔

在循环中加入短暂休眠,避免连续请求触发网站的频率限制:

import time

# 在循环内的合适位置添加,比如driver.get之前
time.sleep(2)

3. 优化等待策略

适当延长WebDriverWait的超时时间,避免因Colab网络延迟导致的假超时:

# 把原有的等待时间从2、4秒适当调整,比如:
wait = WebDriverWait(driver, 5)

核心原因

Colab的服务器IP属于共享资源,CMS网站可能对这类IP的连续请求有更严格的反爬规则;同时复用浏览器会话会携带之前的访问痕迹,进一步触发拦截。独立实例+反检测参数能有效降低被识别为爬虫的概率,缓解超时问题。

内容的提问来源于stack exchange,提问作者Marshall K

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 07:57:33