使用Python爬取网站时如何绕过反爬检测避免被识别为Bot?
绕过coches.net反爬的解决方案
你遇到的403报错和Selenium被检测问题,是因为站点启用了Cloudflare反爬机制,原生Selenium和requests的机器人特征很容易被识别,可按以下方案解决:
方案1:使用undetected-chromedriver绕过Selenium检测
这是落地成本最低的方案,该库会自动抹除Selenium的所有机器人特征,适配Cloudflare的反爬检测:
- 第一步安装依赖:
pip install undetected-chromedriver beautifulsoup4 lxml - 示例代码:
import undetected_chromedriver as uc from bs4 import BeautifulSoup import time # 配置启动参数 options = uc.ChromeOptions() # 模拟正常浏览器的语言设置 options.add_argument("--lang=es-ES") # 固定窗口大小,避免无头特征 options.add_argument("--window-size=1920,1080") driver = uc.Chrome(options=options) URL = 'https://www.coches.net/segunda-mano/' driver.get(URL) # 等待Cloudflare验证完成 time.sleep(5) # 此时可以正常获取页面内容 print(driver.page_source) soup = BeautifulSoup(driver.page_source, 'lxml') # 用完关闭driver driver.quit()
方案2:修复requests请求问题
你之前的请求有两处明显错误:
- 请求头键名错误,正确的UA头键是
User-Agent而非UserAgent - 仅携带UA头不够,需要补充浏览器默认的完整请求头,同时要处理TLS指纹问题
- 推荐替换为httpx库模拟浏览器TLS指纹,示例代码:
import httpx from fake_useragent import UserAgent import random import time ua = UserAgent() headers = { "User-Agent": ua.random, "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8", "Accept-Language": "es-ES,es;q=0.8,en-US;q=0.5,en;q=0.3", "Accept-Encoding": "gzip, deflate, br", "Connection": "keep-alive", "Upgrade-Insecure-Requests": "1", "Sec-Fetch-Dest": "document", "Sec-Fetch-Mode": "navigate", "Sec-Fetch-Site": "none", "Sec-Fetch-User": "?1" } URL = 'https://www.coches.net/segunda-mano/' with httpx.Client(http2=True) as client: r = client.get(URL, headers=headers) print(r.status_code) print(r.text) # 每次请求后加随机等待 time.sleep(random.uniform(1,3))
额外注意事项
- 不要短时间内发起大量请求,每次请求间隔添加1-3秒的随机等待,模拟真人操作
- 如果频繁被拦截,可以绑定动态代理IP,避免IP被封禁
- 尽量不要使用Selenium的无头模式,无头模式下的浏览器特征更容易被反爬系统识别
内容的提问来源于stack exchange,提问作者Georg Klippenstein
相关产品推荐
相关产品推荐

