Google Colab中Selenium访问部分URL超时,顺序影响执行结果求助
解决CMS爬虫在Google Colab中的TimeoutException问题
问题分析
你的爬虫在本地运行正常,但在Colab中连续访问第二个URL时触发driver.get(url)超时,单独访问该URL却能成功。这大概率是因为Colab共享环境下,同一个Chrome会话的缓存、Cookie或连续请求触发了CMS网站的反爬机制,导致后续请求被拦截。
解决方案
1. 为每个URL创建独立的浏览器实例
复用同一个浏览器会话可能会携带之前的会话信息,被目标网站识别为异常请求。修改代码,在循环内初始化新的Chrome实例,每个URL使用独立会话:
from selenium import webdriver from selenium.common.exceptions import TimeoutException from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.wait import WebDriverWait def download_documents() -> None: """Download billing code documents from CMS""" chrome_options = Options() chrome_options.add_argument("--headless") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") # 添加反自动化检测参数 chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) working_url = "https://www.cms.gov/medicare-coverage-database/view/article.aspx?articleid=59626&ver=6" not_working_url = "https://www.cms.gov/medicare-coverage-database/view/lcd.aspx?lcdid=36377&ver=19" for row in [working_url, not_working_url]: print(f"Retrieving from {row}...") # 每个URL新建浏览器实例 driver = webdriver.Chrome(options=chrome_options) try: driver.get(row) print("Wait for webdriver...") wait = WebDriverWait(driver, 2) print("Attempting license accept...") try: wait.until(EC.element_to_be_clickable((By.ID, "btnAcceptLicense"))).click() except TimeoutException: pass wait = WebDriverWait(driver, 4) print("Attempting pop up close...") try: wait.until( EC.element_to_be_clickable( (By.XPATH, "//button[@data-page-action='Clicked the Tracking Sheet Close button.']") ) ).click() except TimeoutException: pass print("Attempting download...") driver.find_element(By.ID, "btnDownload").click() finally: # 确保每个实例使用后关闭 driver.quit() download_documents()
2. 添加请求间隔
在循环中加入短暂休眠,避免连续请求触发网站的频率限制:
import time # 在循环内的合适位置添加,比如driver.get之前 time.sleep(2)
3. 优化等待策略
适当延长WebDriverWait的超时时间,避免因Colab网络延迟导致的假超时:
# 把原有的等待时间从2、4秒适当调整,比如: wait = WebDriverWait(driver, 5)
核心原因
Colab的服务器IP属于共享资源,CMS网站可能对这类IP的连续请求有更严格的反爬规则;同时复用浏览器会话会携带之前的访问痕迹,进一步触发拦截。独立实例+反检测参数能有效降低被识别为爬虫的概率,缓解超时问题。
内容的提问来源于stack exchange,提问作者Marshall K
相关产品推荐
相关产品推荐

