Selenium批量爬取链接报错:目标机器主动拒绝连接
解决方案:Selenium批量爬取连接拒绝错误修复
核心原因排查与修复步骤
1. 检查scrape_info函数是否误关Driver
单独测试正常但批量运行报错,最大可能是scrape_info内部调用了driver.quit()或driver.close(),导致后续循环中Driver会话已失效。
- 修复:打开
scrape_info函数,删除所有关闭Driver的语句,仅保留批量函数finally块中的统一关闭逻辑。
2. 补全Driver初始化参数+添加重试机制
原代码创建Driver时未传入service参数,且缺少异常重试逻辑,容易引发会话中断:
from selenium.common.exceptions import WebDriverException import time def scrape_info_for_pages(links_dict): all_info = {} service = webdriver.chrome.service.Service('C:/Users/...') options = webdriver.ChromeOptions() options.headless = False # 添加稳定性参数 options.add_argument("--no-sandbox") options.add_argument("--disable-dev-shm-usage") options.add_argument("--disable-gpu") # 必须传入service参数,否则Driver启动可能异常 driver = webdriver.Chrome(service=service, options=options) driver.set_page_load_timeout(30) # 设置页面加载超时 try: for key, links in links_dict.items(): info_for_key = [] for url in links: max_retries = 3 retry_count = 0 success = False while retry_count < max_retries and not success: try: result = scrape_info(driver, url) if result: info_for_key.append(result) success = True time.sleep(1) # 添加请求间隔,避免资源过载 except WebDriverException as e: retry_count += 1 print(f"重试第{retry_count}次处理{url},错误:{str(e)}") time.sleep(2) # 会话中断时重建Driver if "Failed to establish a new connection" in str(e): driver.quit() driver = webdriver.Chrome(service=service, options=options) if not success: print(f"多次重试后仍无法处理{url}") all_info[key] = info_for_key finally: driver.quit() return all_info
3. 确认ChromeDriver与浏览器版本严格匹配
即使更新了浏览器,也要保证ChromeDriver主版本号与Chrome完全一致(如Chrome 118对应ChromeDriver 118.x):
- 查看Chrome版本:地址栏输入
chrome://version/ - 下载对应版本的ChromeDriver,确保版本完全匹配
4. 验证流程
- 先测试少量链接(如仅处理Page1的第一个链接),确认无错误
- 逐步增加链接数量,观察运行稳定性
- 若仍有问题,在
scrape_info中添加日志,记录每个步骤的Driver状态
内容的提问来源于stack exchange,提问作者Nico Flor
相关产品推荐
相关产品推荐

