使用Selenium手动解Cloudflare hCaptcha后仍遭拦截,求问题原因与解法
问题:Selenium+stealth仍被Cloudflare识别,手动解hCaptcha后反复出现403
我想搭建一套半自动化爬虫方案,抓取受Cloudflare hCaptcha保护的网站,计划验证码出现时手动解决,之后让爬虫持续抓取到下一次验证码出现。我用Selenium模拟普通用户访问,代码如下:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium_stealth import stealth options = webdriver.ChromeOptions() options.add_argument("start-maximized") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) s=Service(ChromeDriverManager().install()) driver = webdriver.Chrome(service=s, options=options) stealth(driver, languages=["en-US", "en"], vendor="Google Inc.", platform="Win32", webgl_vendor="Intel Inc.", renderer="Intel Iris OpenGL Engine", fix_hairline=True, ) driver.get(url_to_scrape) # Fill the captcha manually
但解决验证码后,Cloudflare还是不让进,页面刷新又显示验证码(返回403),得反复验证。我觉得是Selenium仍被识别成机器人,但已经用了stealth插件伪装,请问哪里操作错了?
可能的问题及解决方法
- 指纹伪装不匹配真实环境:stealth里的
webgl_vendor和renderer是固定值,和你本地真实Chrome环境不符。打开真实Chrome访问about:gpu,查看自己的WebGL供应商和渲染器,替换代码里的对应参数,确保指纹完全匹配。 - 缺少真实用户行为模拟:Cloudflare会检测鼠标移动、页面滚动等行为,你现在直接打开页面就解验证码,太像机器人。可以在
driver.get(url)后添加模拟操作:比如用ActionChains缓慢移动鼠标到验证码区域、随机滚动页面,再等待手动验证。 - Chrome驱动与浏览器版本不兼容:WebDriverManager自动安装的驱动可能和本地Chrome版本有细微差异,触发识别。手动下载和你Chrome版本完全一致的ChromeDriver,替换自动安装的路径。
- 额外浏览器特征暴露:补充这些参数强化伪装:
options.add_argument("--disable-autofill")禁用自动填充options.add_argument("--disable-extensions")禁用扩展options.add_argument("--disable-dev-shm-usage")关闭开发者工具相关特性options.add_argument("--no-sandbox")禁用沙箱模式options.add_argument("--disable-blink-features=AutomationControlled")屏蔽自动化控制标记
- 会话复用与Cookie问题:手动解完验证码后,必须复用同一个driver实例,不能重新创建。Cloudflare的验证Cookie有有效期,确保爬虫后续请求都在当前会话内进行。
- 请求频率过高触发二次验证:即使通过验证,短时间高频请求也会被拦截。在抓取逻辑中添加随机等待时间(比如
time.sleep(random.uniform(2,5))),模拟真实用户浏览间隔。
优化后的示例代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.common.by import By from selenium_stealth import stealth import time import random options = webdriver.ChromeOptions() # 基础伪装参数 options.add_argument("start-maximized") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) # 补充特征隐藏参数 options.add_argument("--disable-autofill") options.add_argument("--disable-extensions") options.add_argument("--disable-dev-shm-usage") options.add_argument("--no-sandbox") options.add_argument("--disable-blink-features=AutomationControlled") # 建议手动指定匹配本地Chrome版本的Driver路径,替换为你的实际路径 # s = Service("C:/path/to/your/chromedriver.exe") s = Service(ChromeDriverManager().install()) driver = webdriver.Chrome(service=s, options=options) # 替换为你真实Chrome的参数(从about:gpu获取) stealth(driver, languages=["zh-CN", "zh"], vendor="Google Inc.", platform="Win32", webgl_vendor="NVIDIA Corporation", renderer="NVIDIA GeForce GTX 1650/PCIe/SSE2", fix_hairline=True, run_on_insecure_origins=True, ) driver.get(url_to_scrape) # 模拟真实用户操作路径 time.sleep(random.uniform(1,3)) driver.execute_script("window.scrollTo(0, document.body.scrollHeight/2);") time.sleep(random.uniform(1,2)) # 定位验证码区域并移动鼠标 captcha_iframe = driver.find_element(By.CSS_SELECTOR, 'iframe[src*="hcaptcha"]') ActionChains(driver).move_to_element(captcha_iframe).perform() time.sleep(random.uniform(0.5,1.5)) print("请手动完成验证码验证,完成后按回车继续...") input() # 持续抓取逻辑,复用会话并检测验证码 while True: # 你的内容抓取代码 print("正在抓取当前页面内容...") # 模拟用户浏览间隔 time.sleep(random.uniform(3,6)) # 检查是否再次出现验证码 try: driver.find_element(By.CSS_SELECTOR, 'iframe[src*="hcaptcha"]') print("再次触发验证码,请手动验证后按回车...") input() except: continue
内容的提问来源于stack exchange,提问作者Laszer
相关产品推荐
相关产品推荐

