Python/Selenium爬取谷歌页面遇MaxRetryError[Errno61]连接拒绝求助
问题分析
你遇到的MaxRetryError: [Errno 61] Connection refused核心原因有两个:一是你在chrome.quit()后没有重新初始化ChromeDriver实例(代码里注释掉了chrome = webdriver.Chrome()),导致后续调用chrome.get()时,试图连接已经被销毁的会话;二是谷歌的反爬机制可能已经识别到你的自动化请求,拒绝了连接。
1. 具体解决方法
(1)修复Driver实例重建逻辑
这是当前最关键的问题!你在调用chrome.quit()关闭驱动后,必须重新创建一个新的ChromeDriver实例,否则后续的chrome.get()会尝试连接已失效的会话,必然触发连接拒绝错误。修改你的代码块:
domain = pattern.search(website) counter = 2 # keep running this until the url appears like normal while domain is None: counter += 1 # close chrome and try again print('link not found, closing chrome and restarting ...\nwaiting {} seconds...'.format(counter)) chrome.quit() time.sleep(counter) # 重新初始化ChromeDriver实例(取消注释并确保配置正确) chrome = webdriver.Chrome() # 这里可以添加ChromeOptions配置增强反爬能力 time.sleep(10) chrome.get('https://google.com') # ... 后续代码保持不变
(2)增强反爬规避策略
谷歌对自动化工具的检测很严格,仅靠sleep不够,建议添加以下配置:
- 禁用自动化特征检测:
from selenium.webdriver.chrome.options import Options chrome_options = Options() # 禁用Chrome的自动化提示和特征 chrome_options.add_argument('--disable-blink-features=AutomationControlled') chrome_options.add_argument('--disable-dev-shm-usage') chrome_options.add_argument('--no-sandbox') # 设置真实的User-Agent,模拟普通浏览器 chrome_options.add_argument('user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36') # 初始化driver时传入配置 chrome = webdriver.Chrome(options=chrome_options)
- 使用随机延迟替代固定计数器:避免请求间隔有规律,比如用
random.randint(5, 15)代替固定的counter值。 - 使用显式等待替代固定
sleep:等待元素加载完成后再操作,更可靠且节省时间:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 等待搜索框出现后再输入 target = WebDriverWait(chrome, 10).until( EC.presence_of_element_located((By.NAME, 'q')) ) target.send_keys(college) target.send_keys(Keys.RETURN)
(3)确保Driver生命周期正确管理
如果你的脚本会被循环调用,建议每次调用函数时都重新初始化Driver,并用try-finally块确保无论是否出错都能关闭Driver:
def your_scraping_function(college): chrome = None try: chrome_options = Options() # 添加上述配置 chrome = webdriver.Chrome(options=chrome_options) # ... 你的爬取逻辑 except Exception as e: print(f"Error occurred: {e}") finally: if chrome is not None: chrome.quit()
2. 在Selenium上下文捕获该错误
你可以通过try-except块捕获urllib3.exceptions.MaxRetryError,同时也建议捕获Selenium的WebDriverException(因为连接错误通常会被包装成这个异常):
from urllib3.exceptions import MaxRetryError from selenium.common.exceptions import WebDriverException try: chrome.get('https://google.com') except (MaxRetryError, WebDriverException) as e: print(f"Connection failed: {e}") # 这里可以添加重试逻辑,比如重新初始化Driver后再次尝试 chrome.quit() chrome = webdriver.Chrome(options=chrome_options) chrome.get('https://google.com')
如果需要更精细的捕获,可以检查异常的args来判断是否是连接拒绝错误:
try: chrome.get('https://google.com') except WebDriverException as e: if "Connection refused" in str(e): print("Connection was refused by Google, retrying...") # 执行重试逻辑 chrome.quit() chrome = webdriver.Chrome(options=chrome_options) chrome.get('https://google.com') else: # 处理其他Selenium错误 raise e
内容的提问来源于stack exchange,提问作者im2wddrf
相关产品推荐
相关产品推荐

