Selenium driver.get()仅在特定网站卡顿问题排查求助
问题:访问Student Beans特定页面时Selenium ChromeDriver长时间卡顿
问题概述
我是网页爬取新手,该问题困扰我已久却无法解决。我需要爬取列表中的多个网站,Chrome Driver对其他网站均正常,但访问https://www.studentbeans.com/student-discount/it/cats时,driver.get()请求会长时间卡顿(有时超1小时),必须手动中断。
该代码数月前运行正常,如今仅该链接出现问题。我已查阅资料并添加各类Driver选项,但均无效。
版本信息
- Chrome Driver版本:107.0.5304.62(已更新)
- Chrome版本:107.0.5304.107(已更新)
- Selenium版本:4.2.0
Driver连接器代码
class SeleniumConnector() : def __init__(self) -> None: self.options = webdriver.ChromeOptions() def connector(self, link: str): """Creates a driver instance for google Chrome and allows to connect to given link""" self.options.add_argument('--ignore-certificate-errors') self.options.add_argument("--disable-notifications") #disabling notifications self.options.add_argument("--disable-popup-blocking") #to avoid chat popus self.options.add_argument('--disable-gpu') if os.name == 'nt' else None self.options.add_argument("--incognito") self.options.add_argument("--disable-blink-features") self.options.add_argument("--disable-blink-features=AutomationControlled") self.options.add_argument("--headless") #to avoid popups of every window desired_capabilities = DesiredCapabilities.CHROME.copy() desired_capabilities['acceptInsecureCerts'] = True self.options.add_experimental_option("excludeSwitches", ["enable-automation"]) self.options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options= self.options) driver.execute_cdp_cmd('Network.setUserAgentOverride', {"userAgent": 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/83.0.4103.53 Safari/537.36'}) try: print('Connecting to..... {}'.format(link)) driver.delete_all_cookies() driver.set_page_load_timeout(15) driver.get(link) except: print(f'Wrong link: {link}') return driver
键盘中断后的回溯信息
[WDM] - ====== WebDriver manager ====== 2022-11-12 13:48:17,414 INFO ====== WebDriver manager ====== return self.request_encode_body( File "C:\Users\---\anaconda3\lib\site-packages\urllib3\request.py", line 170, in request_encode_body return self.urlopen(method, url, **extra_kw) File "C:\Users\---\anaconda3\lib\site-packages\urllib3\poolmanager.py", line 376, in urlopen response = conn.urlopen(method, u.request_uri, **kw) File "C:\Users\---\anaconda3\lib\site-packages\urllib3\connectionpool.py", line 703, in urlopen httplib_response = self._make_request( File "C:\Users\---\anaconda3\lib\site-packages\urllib3\connectionpool.py", line 449, in _make_request six.raise_from(e, None) File "<string>", line 3, in raise_from File "C:\Users\---\anaconda3\lib\site-packages\urllib3\connectionpool.py", line 444, in _make_request httplib_response = conn.getresponse() File "C:\Users\---\anaconda3\lib\http\client.py", line 1377, in getresponse response.begin() File "C:\Users\---\anaconda3\lib\http\client.py", line 320, in begin version, status, reason = self._read_status() File "C:\Users\fetza\anaconda3\lib\http\client.py", line 281, in _read_status line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1") File "C:\Users\---\anaconda3\lib\socket.py", line 704, in readinto return self._sock.recv_into(b) KeyboardInterrupt
原因分析
- 反爬机制升级:Student Beans大概率更新了反爬策略,针对无头浏览器、自动化特征做了更严格的检测。现有代码的反检测选项不足以规避识别,导致服务器故意延迟响应或不返回内容。
- 页面资源加载阻塞:该页面包含大量第三方资源(广告、追踪脚本),无头模式下这些资源加载异常,导致页面一直处于加载状态,而现有异常处理未正确捕获超时问题。
- 版本细微不匹配:Chrome版本(107.0.5304.107)和ChromeDriver版本(107.0.5304.62)存在小版本差异,可能引发通信兼容性问题。
- 用户代理不一致:设置的UA是Chrome 83版本,与实际使用的Chrome 107不符,容易被识别为自动化工具。
解决方法
1. 修复版本匹配问题
直接下载与Chrome版本完全一致的ChromeDriver,替换WebDriverManager自动安装的版本,确保版本号完全对应。
2. 优化ChromeOptions增强反检测
在现有选项基础上添加以下配置:
# 模拟真实用户环境 self.options.add_argument("--start-maximized") self.options.add_argument("--disable-dev-shm-usage") self.options.add_argument("--no-sandbox") self.options.add_argument("--disable-extensions") # 设置页面加载策略为eager,仅等待DOM加载完成,不等待所有资源 self.options.page_load_strategy = 'eager' # 无头模式额外伪装(若保留无头) self.options.add_argument("--window-size=1920,1080")
3. 修正异常处理逻辑
区分超时异常与其他错误,超时后主动停止页面加载:
from selenium.common.exceptions import TimeoutException try: print('Connecting to..... {}'.format(link)) driver.delete_all_cookies() driver.set_page_load_timeout(15) driver.get(link) except TimeoutException: print(f'Page load timed out: {link}') driver.execute_script("window.stop();") except Exception as e: print(f'Error connecting to {link}: {str(e)}')
4. 使用匹配版本的用户代理
替换UA为当前Chrome版本对应的字符串:
driver.execute_cdp_cmd('Network.setUserAgentOverride', { "userAgent": 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36' })
5. 尝试非无头模式测试
暂时注释掉--headless选项,验证非无头模式是否能正常加载。若可以,说明无头模式被针对性检测,可尝试升级Chrome到112+版本并使用--headless=new参数。
6. 添加显式等待机制
等待页面关键元素加载替代单纯的页面加载超时:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By try: driver.get(link) WebDriverWait(driver, 10).until( EC.title_contains("Student Discounts") ) except TimeoutException: print(f'Page load or element wait timed out: {link}') driver.execute_script("window.stop();")
内容的提问来源于stack exchange,提问作者afzde
相关产品推荐
相关产品推荐

