Colab中Selenium Chrome意外退出的修复及真人浏览器模拟实现
Colab中Selenium Headless Chrome启动失败(SessionNotCreatedException)解决方法
问题背景
需要爬取https://clutch.co/il/it-services数据,目标网站有CloudFlare反爬机制,计划用Selenium搭配Headless Chrome模拟真人浏览器,但在Google Colab运行时触发SessionNotCreatedException,Chrome启动失败,报错核心信息:
session not created: Chrome failed to start: exited normally.
(session not created: DevToolsActivePort file doesn't exist)
报错原因
Colab是无图形界面的Linux环境,默认的Headless Chrome配置缺少适配Colab环境的关键参数:
- Colab以root用户运行,Chrome默认禁止root使用沙箱模式
- 共享内存空间不足导致启动失败
- 旧版Headless模式参数无法适配Colab环境
- 缺少模拟真实浏览器的必要配置,容易被反爬机制识别
解决方案
通过调整ChromeOptions参数,适配Colab环境并模拟真人浏览器:
步骤1:更新依赖
在Colab中先执行以下命令,确保Selenium版本为最新:
!pip install -U selenium
步骤2:配置正确的ChromeOptions
添加以下关键参数,覆盖原代码中的Options配置:
--headless=new:启用新版无头模式,行为更接近正常Chrome--no-sandbox:禁用沙箱模式,适配root用户运行环境--disable-dev-shm-usage:解决Colab共享内存不足问题--window-size=1920,1080:指定窗口大小,避免页面布局异常- 自定义
user-agent:模拟真实浏览器请求头,绕过反爬检测 --disable-blink-features=AutomationControlled:禁止Chrome暴露自动化工具标识
完整可运行代码
import pandas as pd from bs4 import BeautifulSoup from tabulate import tabulate from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager # 配置Chrome选项 options = Options() # 新版无头模式 options.add_argument("--headless=new") # 适配Colab环境的必要参数 options.add_argument("--no-sandbox") options.add_argument("--disable-dev-shm-usage") options.add_argument("--window-size=1920,1080") # 模拟真实浏览器UA options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36") # 避免被检测为自动化工具 options.add_argument("--disable-blink-features=AutomationControlled") # 禁用GPU加速(可选) options.add_argument("--disable-gpu") # 自动管理ChromeDriver版本,避免版本不匹配 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) url = "https://clutch.co/il/it-services" driver.get(url) html = driver.page_source soup = BeautifulSoup(html, 'html.parser') # 爬取数据逻辑 company_names = soup.select(".directory-list div.provider-info--header .company_info a") locations = soup.select(".locality") company_names_list = [name.get_text(strip=True) for name in company_names] locations_list = [location.get_text(strip=True) for location in locations] data = {"Company Name": company_names_list, "Location": locations_list} df = pd.DataFrame(data) df.index += 1 print(tabulate(df, headers="keys", tablefmt="psql")) df.to_csv("it_services_data.csv", index=False) driver.quit()
额外说明
- 使用
webdriver_manager自动匹配ChromeDriver版本,避免手动下载的版本不兼容问题 - 若仍触发CloudFlare检测,可添加隐式等待
driver.implicitly_wait(10),或模拟页面滚动等真人交互行为
内容的提问来源于stack exchange,提问作者zero
相关产品推荐
相关产品推荐

