Python爬取URL获取HTML和图片遇HTTP 403及Cloudflare验证码问题
解决方案
方案1:使用undetected-chromedriver绕过Cloudflare验证(最稳定)
原生Selenium存在大量可被Cloudflare识别的自动化特征,导致频繁触发验证码。undetected-chromedriver会自动抹除浏览器指纹,验证通过率远高于普通Selenium,同时支持直接获取原始图片资源无需截图。
安装依赖
pip install undetected-chromedriver requests
代码示例(同时支持HTML获取+原图下载)
import undetected_chromedriver as uc import requests import time # 浏览器配置 options = uc.ChromeOptions() # 如需要后台运行可开启无头模式 # options.add_argument("--headless=new") options.add_argument("--start-maximized") options.add_argument("--disable-blink-features=AutomationControlled") driver = uc.Chrome(options=options) # 1. 获取页面HTML内容 target_url = "https://bama.ir/car/detail-grxi644n-hyundai-genesis-coupe-2011" driver.get(target_url) # 等待页面加载、Cloudflare验证完成,可根据网络情况调整等待时长 time.sleep(3) html_content = driver.page_source print(html_content) # 2. 下载原始图片资源(无需截图) # 复用浏览器会话的Cookies和UA,避免CDN返回403 cookies = {item["name"]: item["value"] for item in driver.get_cookies()} headers = { "User-Agent": driver.execute_script("return navigator.userAgent"), "Referer": "https://bama.ir/" } img_url = "https://cdn.bama.ir/uploads/BamaImages/VehicleCarImages/a36114cd-1978-41a4-a558-cbe5f652faf1/CarImage9473771_0_1_thumb_900_600.jpg" img_response = requests.get(img_url, headers=headers, cookies=cookies) with open(img_url.rsplit("/",1)[1], "wb") as f: f.write(img_response.content) driver.quit()
方案2:Playwright+Stealth插件(备选方案)
如果遇到小概率验证失败的情况,可以换用伪装度更高的Playwright框架搭配Stealth插件使用:
安装依赖
pip install playwright playwright-stealth playwright install chrome
关键注意事项
- 控制请求间隔在3-5秒每次,短时间高频请求会触发Cloudflare流量异常判定,强制弹出验证码
- 若本地IP频繁触发验证,可更换住宅代理IP,数据中心IP被Cloudflare标记的概率远高于住宅IP
- 访问页面和下载图片的会话信息必须一致,否则CDN仍会返回403错误
内容的提问来源于stack exchange,提问作者Ali Gh
相关产品推荐
相关产品推荐

