You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Colab中Selenium Chrome意外退出的修复及真人浏览器模拟实现

Colab中Selenium Headless Chrome启动失败(SessionNotCreatedException)解决方法

问题背景

需要爬取https://clutch.co/il/it-services数据,目标网站有CloudFlare反爬机制,计划用Selenium搭配Headless Chrome模拟真人浏览器,但在Google Colab运行时触发SessionNotCreatedException,Chrome启动失败,报错核心信息:

session not created: Chrome failed to start: exited normally.
(session not created: DevToolsActivePort file doesn't exist)

报错原因

Colab是无图形界面的Linux环境,默认的Headless Chrome配置缺少适配Colab环境的关键参数:

  • Colab以root用户运行,Chrome默认禁止root使用沙箱模式
  • 共享内存空间不足导致启动失败
  • 旧版Headless模式参数无法适配Colab环境
  • 缺少模拟真实浏览器的必要配置,容易被反爬机制识别

解决方案

通过调整ChromeOptions参数,适配Colab环境并模拟真人浏览器:

步骤1:更新依赖

在Colab中先执行以下命令,确保Selenium版本为最新:

!pip install -U selenium

步骤2:配置正确的ChromeOptions

添加以下关键参数,覆盖原代码中的Options配置:

  • --headless=new:启用新版无头模式,行为更接近正常Chrome
  • --no-sandbox:禁用沙箱模式,适配root用户运行环境
  • --disable-dev-shm-usage:解决Colab共享内存不足问题
  • --window-size=1920,1080:指定窗口大小,避免页面布局异常
  • 自定义user-agent:模拟真实浏览器请求头,绕过反爬检测
  • --disable-blink-features=AutomationControlled:禁止Chrome暴露自动化工具标识

完整可运行代码

import pandas as pd
from bs4 import BeautifulSoup
from tabulate import tabulate
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager

# 配置Chrome选项
options = Options()
# 新版无头模式
options.add_argument("--headless=new")
# 适配Colab环境的必要参数
options.add_argument("--no-sandbox")
options.add_argument("--disable-dev-shm-usage")
options.add_argument("--window-size=1920,1080")
# 模拟真实浏览器UA
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36")
# 避免被检测为自动化工具
options.add_argument("--disable-blink-features=AutomationControlled")
# 禁用GPU加速(可选)
options.add_argument("--disable-gpu")

# 自动管理ChromeDriver版本,避免版本不匹配
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)

url = "https://clutch.co/il/it-services"
driver.get(url)

html = driver.page_source
soup = BeautifulSoup(html, 'html.parser')

# 爬取数据逻辑
company_names = soup.select(".directory-list div.provider-info--header .company_info a")
locations = soup.select(".locality")

company_names_list = [name.get_text(strip=True) for name in company_names]
locations_list = [location.get_text(strip=True) for location in locations]

data = {"Company Name": company_names_list, "Location": locations_list}
df = pd.DataFrame(data)
df.index += 1
print(tabulate(df, headers="keys", tablefmt="psql"))
df.to_csv("it_services_data.csv", index=False)

driver.quit()

额外说明

  • 使用webdriver_manager自动匹配ChromeDriver版本,避免手动下载的版本不兼容问题
  • 若仍触发CloudFlare检测,可添加隐式等待driver.implicitly_wait(10),或模拟页面滚动等真人交互行为

内容的提问来源于stack exchange,提问作者zero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 23:48:11