Scrapy+Scrapy-Selenium出现WinError 10061错误的原因与解决咨询
WinError 10061错误排查与修复(Scrapy+Scrapy-Selenium爬虫)
错误现象
基于Scrapy+Scrapy-Selenium开发的爬虫,在成功爬取一段时间后出现WinError 10061错误,具体日志如下:
2023-02-22 06:12:38 [selenium.webdriver.remote.remote_connection] DEBUG: DELETE http://localhost:58736/session/ec60bfa108737c8be99566052d54e5d3 {} 2023-02-22 06:12:38 [urllib3.connectionpool] DEBUG: Resetting dropped connection: localhost 2023-02-22 06:12:42 [urllib3.util.retry] DEBUG: Incremented Retry for (url='/session/ec60bfa108737c8be99566052d54e5d3'): Retry(total=2, connect=None, read=None, redirect=None, status=None) 2023-02-22 06:12:42 [urllib3.connectionpool] WARNING: Retrying (Retry(total=2, connect=None, read=None, redirect=None, status=None)) after connection broken by 'NewConnectionError(<urllib3.connection.HTTPConnection object at 0x000001FA6D516BE0>: Failed to establish a new connection: [WinError 10061] No connection could be made because the target machine actively refused it)': /session/ec60bfa108737c8be99566052d54e5d3 2023-02-22 06:12:42 [urllib3.connectionpool] DEBUG: Starting new HTTP connection (2): localhost:58736 2023-02-22 06:12:46 [urllib3.util.retry] DEBUG: Incremented Retry for (url='/session/ec60bfa108737c8be99566052d54e5d3'): Retry(total=1, connect=None, read=None, redirect=None, status=None) 2023-02-22 06:12:46 [urllib3.connectionpool] WARNING: Retrying (Retry(total=1, connect=None, read=None, redirect=None, status=None)) after connection broken by 'NewConnectionError(<urllib3.connection.HTTPConnection object at 0x000001FA6DAE3D00>: Failed to establish a new connection: [WinError 10061] No connection could be made because the target machine actively refused it)': /session/ec60bfa108737c8be99566052d54e5d3 2023-02-22 06:12:46 [urllib3.connectionpool] DEBUG: Starting new HTTP connection (3): localhost:58736 2023-02-22 06:12:50 [urllib3.util.retry] DEBUG: Incremented Retry for (url='/session/ec60bfa108737c8be99566052d54e5d3'): Retry(total=0, connect=None, read=None, redirect=None, status=None) 2023-02-22 06:12:50 [urllib3.connectionpool] WARNING: Retrying (Retry(total=0, connect=None, read=None, redirect=None, status=None)) after connection broken by 'NewConnectionError(<urllib3.connection.HTTPConnection object at 0x000001FA6D63BA30>: Failed to establish a new connection: [WinError 10061] No connection could be made because the target machine actively refused it)': /session/ec60bfa108737c8be99566052d54e5d3 2023-02-22 06:12:50 [urllib3.connectionpool] DEBUG: Starting new HTTP connection (4): localhost:58736
产生原因
WinError 10061本质是本地无法连接到Selenium WebDriver服务,结合你的爬虫代码,主要原因包括:
- 浏览器驱动进程长时间复用崩溃:
click方法持续复用同一个浏览器实例,循环爬取多页后,浏览器内存泄漏、进程卡死或被系统强制回收,导致Scrapy-Selenium无法与驱动建立连接。 - 固定
time.sleep加剧资源消耗:大量使用time.sleep()等待页面加载,不仅浪费资源,还可能因页面未完全加载就执行操作引发异常,加速驱动进程崩溃。 - 未主动释放驱动资源:循环结束或出现异常时,未调用
browser.quit()关闭浏览器,导致驱动进程后台残留,后续请求仍尝试连接已失效的会话。
修复方法
1. 替换固定sleep为Selenium显式等待
使用显式等待替代time.sleep(),精准等待元素加载完成,减少不必要的等待时间,降低资源占用:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 示例:等待下拉框可点击 WebDriverWait(browser, 10).until( EC.element_to_be_clickable((By.XPATH, "/html/body/div[1]/div/section[2]/div/div/div/div[1]/div[2]/select")) ).click()
2. 主动管理浏览器生命周期
在循环结束或异常时,主动关闭浏览器并释放资源;如果爬取页数较多,可设置每爬取N页后重启浏览器:
def click(self, response): browser = response.meta['driver'] try: # 原有的下拉选择、循环爬取逻辑... next_page = True page_count = 0 while next_page: page_count +=1 # 爬取逻辑... # 每5页重启一次浏览器 if page_count %5 ==0: browser.quit() # 重新发起SeleniumRequest获取新的浏览器实例 yield SeleniumRequest(url=browser.current_url, wait_time=5, callback=self.click) return # 下一页逻辑... except Exception as e: print(f"Error occurred: {e}") finally: # 无论成功失败,都关闭浏览器 browser.quit()
3. 配置Scrapy-Selenium的会话超时
在Scrapy的settings.py中配置Selenium驱动的超时参数,避免长时间等待失效连接:
SELENIUM_DRIVER_ARGUMENTS = ['--headless=new', '--disable-gpu', '--no-sandbox'] SELENIUM_TIMEOUT = 30 SELENIUM_REQUEST_DELAY = 1
4. 优化爬虫请求逻辑
如果tournament详情页需要JS渲染,将scrapy.Request改为SeleniumRequest;若无需JS渲染,直接用普通请求即可,避免混用请求类型导致会话冲突。
代码优化示例(核心部分)
修改click方法,加入显式等待和资源回收:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC def click(self, response): browser = response.meta['driver'] try: # 等待第一个下拉框并点击 year_select = WebDriverWait(browser, 10).until( EC.element_to_be_clickable((By.XPATH, "/html/body/div[1]/div/section[2]/div/div/div/div[1]/div[2]/select")) ) year_select.click() # 等待年份选项并点击 year_option = WebDriverWait(browser, 10).until( EC.element_to_be_clickable((By.XPATH, '//*[@id="search_year"]/option[6]')) ) year_option.click() # 同理处理第二个下拉框 type_select = WebDriverWait(browser, 10).until( EC.element_to_be_clickable((By.XPATH, '/html/body/div[1]/div/section[2]/div/div/div/div[1]/div[1]/select')) ) type_select.click() type_option = WebDriverWait(browser, 10).until( EC.element_to_be_clickable((By.XPATH, "/html/body/div[1]/div/section[2]/div/div/div/div[1]/div[1]/select/option[5]")) ) type_option.click() next_page = True page_count = 0 while next_page: page_count +=1 html = browser.page_source response = Selector(text=html) tournaments_urls = response.xpath('//header[@class="entity_header"]/h2/a/@href') for el in tournaments_urls.extract(): yield scrapy.Request(url=el, callback=self.parse, dont_filter=True) try: next_page_button = WebDriverWait(browser, 10).until( EC.element_to_be_clickable((By.XPATH, "//html/body/div[1]/div/section[2]/div/div/div/div[3]/ul[2]/li[11]/a")) ) next_page_button.click() print("next page!!!") # 等待页面加载完成 WebDriverWait(browser, 10).until( EC.presence_of_element_located((By.XPATH, '//header[@class="entity_header"]/h2/a')) ) except: print("No next page!!!!!!!!!!") next_page = False break # 每5页重启浏览器 if page_count %5 ==0: browser.quit() yield SeleniumRequest(url=browser.current_url, wait_time=5, callback=self.click) return except Exception as e: print(f"Error in click method: {e}") finally: browser.quit()
内容的提问来源于stack exchange,提问作者NoName123576
相关产品推荐
相关产品推荐

