You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium Chrome批量爬取站点时触发InvalidSessionId异常排查

问题背景

基于Selenium编写的URL爬取逻辑接入pandas DataFrame批量处理流程,目标爬取规模为10⁴个站点。代码成功爬取约50个站点后固定抛出InvalidSessionId错误,逻辑中仅在每批爬取末尾显式关闭驱动,运行环境为Google Colab。

单站点爬取实现

功能为加载目标URL、提取页面body文本、捕获WebDriverException异常:

def scrape_all_text(url, keyword, wd):
  try:
    print(str(url))
    if (str(url).startswith("http://") or str(url).startswith("https://")):
      wd.get(str(url))
    else:
      wd.get("http://" + str(url))
    text = wd.find_element_by_tag_name("body").text.replace('\n', ' ')
    print(f"KEYWORD: {keyword}, TEXT: {text}")
    return text
  except WebDriverException as e:
     print(f"KEYWORD: {keyword}, TEXT: {None}, EXCEPTION: {e}")
     return None

分批爬取生成器实现

按比例拆分数据集,每批启动新的Chrome Webdriver实例,设置20秒页面加载超时,逐行调用爬取函数后关闭驱动并返回批次结果:

def split_and_scrape(split_percent, df, col_to_add, scrape_func):
  num_splits = math.ceil(np.reciprocal(split_percent))
  entries_per_split = int(len(df.index) * split_percent)
  split_df_list = np.array_split(df, num_splits)
  for i, split in enumerate(split_df_list):
    wd = webdriver.Chrome('chromedriver',chrome_options=chrome_options)
    wd.set_page_load_timeout(20)
    print(f"Running on {entries_per_split*i}th - {entries_per_split*(i+1)}th entries")
    split[col_to_add] = split.apply(lambda x: scrape_func(x['guess_site_url'], x['keyword'], wd), axis=1)
    wd.close()
    yield split

完整报错栈

InvalidSessionIdException                 Traceback (most recent call last)
<ipython-input-90-55b98c96b157> in <module>()
      2 # wd.set_page_load_timeout(20)
      3 # merged['page_contents'] = merged.apply(lambda x: scrape_all_text(x['guess_site_url'], x['keyword'], wd), axis=1) #next put in function where merged saves every few entries
----> 4 for i, split in enumerate(split_and_scrape(0.001, merged, 'page_contents', scrape_all_text)):
      5   split.to_csv(f"page_contents_{i}.csv")

3 frames
/usr/local/lib/python3.7/dist-packages/selenium/webdriver/remote/errorhandler.py in check_response(self, response)
    245                 alert_text = value['alert'].get('text')
    246             raise exception_class(message, screen, stacktrace, alert_text)  # type: ignore[call-arg]  # mypy is not smart enough here
--> 247         raise exception_class(message, screen, stacktrace)
    248 
    249     def _value_or_default(self, obj: Mapping[_KT, _VT], key: _KT, default: _VT) -> _VT:

InvalidSessionIdException: Message: invalid session id
触发原因
  • 驱动关闭方法用错:wd.close()仅关闭当前浏览器窗口,不会终止ChromeDriver后台进程、也不会彻底释放会话资源。多批次运行后残留进程堆积,占满Colab分配的运行配额,会话直接失效。
  • 无会话存活校验:页面加载超时、站点触发浏览器崩溃、Colab后台回收闲置渲染进程时,会直接杀死WebDriver会话,现有逻辑没有检测会话状态,后续请求仍调用已死亡的驱动实例,直接抛出无效ID错误。
  • 异常捕获不全:现有逻辑仅捕获WebDriverException,没有单独处理会话失效场景,出现问题后不会自动重建驱动,只会持续报错。
修复方案
  • 将每批末尾的wd.close()替换为wd.quit(),每批跑完彻底终止ChromeDriver进程,释放全部会话资源,避免残留进程堆积。
  • 增加会话有效性检测,每次发起爬取请求前先校验驱动是否可用,发现会话失效立刻重建驱动实例。
  • 缩小单批次爬取规模,单个Chrome实例连续爬取10-20个站点就主动重启,避免长时间运行触发Colab的进程回收机制。
  • 扩展异常捕获逻辑,单独识别InvalidSessionId错误,捕获后标记当前会话已损坏,下次请求自动重建驱动。

修复后的分批逻辑参考:

from selenium.common.exceptions import InvalidSessionIdException

def split_and_scrape(split_percent, df, col_to_add, scrape_func):
  num_splits = math.ceil(np.reciprocal(split_percent))
  entries_per_split = int(len(df.index) * split_percent)
  split_df_list = np.array_split(df, num_splits)
  for i, split in enumerate(split_df_list):
    # 初始化驱动
    wd = webdriver.Chrome('chromedriver',chrome_options=chrome_options)
    wd.set_page_load_timeout(20)
    print(f"Running on {entries_per_split*i}th - {entries_per_split*(i+1)}th entries")
    scrape_results = []
    for _, row in split.iterrows():
        try:
            # 先校验会话是否存活
            wd.current_url
        except InvalidSessionIdException:
            # 会话失效则重建
            wd.quit()
            wd = webdriver.Chrome('chromedriver',chrome_options=chrome_options)
            wd.set_page_load_timeout(20)
        res = scrape_func(row['guess_site_url'], row['keyword'], wd)
        scrape_results.append(res)
    split[col_to_add] = scrape_results
    # 用quit彻底关闭驱动
    wd.quit()
    yield split

内容的提问来源于stack exchange,提问作者Omar Dahleh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 12:25:18