如何使用Selenium Webdriver规避沃尔玛网站的访问拦截
解决方案
- 替换原生chromedriver为
undetected-chromedriver
这是目前绕过电商站点自动化检测最高效的方案,它默认修改了所有selenium可被检测的特征指纹,安装和修改后代码示例如下:
首先执行安装命令:pip install undetected-chromedriver
修改后的代码:
import undetected_chromedriver as uc import time # 可以继续保留你原有的用户数据目录配置 options = uc.ChromeOptions() options.add_argument(r"user-data-dir=C:\Users\PCHOME\AppData\Local\Temp\scoped1\Default\gdps") driver = uc.Chrome(driver_executable_path=r'dchrome\chromedriver.exe', options=options) driver.get("https://www.walmart.com/")
大部分情况下完成这一步就不会再反复弹出验证码,可正常登录。
- 原生selenium添加反检测配置
如果不想引入第三方依赖,可以给原生启动参数添加反爬配置,覆盖自动化特征标记:
from selenium import webdriver from selenium.webdriver.chrome.options import Options import time options = Options() # 核心反检测参数 options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option("useAutomationExtension", False) options.add_argument(r"user-data-dir=C:\Users\PCHOME\AppData\Local\Temp\scoped1\Default\gdps") driver = webdriver.Chrome(executable_path=r'dchrome\chromedriver.exe', options=options) # 启动后修改webdriver属性标记 driver.execute_cdp_cmd("Page.addScriptToEvaluateOnNewDocument", { "source": """ Object.defineProperty(navigator, 'webdriver', { get: () => undefined }) """ }) driver.get("https://www.walmart.com/")
- 优化操作行为降低异常标记
- 建议用户数据目录替换为你平时正常使用Chrome的个人数据目录,而非临时目录,正常的浏览历史、Cookie记录会降低反爬系统的异常判定概率
- 所有模拟操作添加随机延迟,不要使用固定时长的sleep,示例:
time.sleep(random.uniform(0.8, 2.5)),点击、输入等操作的间隔符合真人使用习惯 - 不要短时间内频繁刷新页面、重复提交请求
内容的提问来源于stack exchange,提问作者testgen
相关产品推荐
相关产品推荐

