如何修改Selenium爬虫代码规避rome2rio的验证码检测?
爬取rome2rio网站规避验证码的优化方案
我尝试爬取rome2rio网站时,99%的情况都会遇到验证码,以下是我当前使用的代码,求修改或优化建议以规避检测:
当前代码
from selenium import webdriver import undetected_chromedriver as uc import time import random # Initialize undetected ChromeOptions chrome_options = uc.ChromeOptions() # Essential options to avoid detection chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") chrome_options.add_argument("--incognito") # Correctly setting excludeSwitches within undetected_chromedriver context chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_argument("--start-maximized") # To start maximized chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) # Rotating User-Agent user_agents = [ "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36", "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36", # Add more as needed ] random_user_agent = random.choice(user_agents) chrome_options.add_argument(f"user-agent={random_user_agent}") # Adjusting viewport size to non-standard dimensions if needed # chrome_options.add_argument("--window-size=1366,768") # Use only if you don't want to start maximized # Use undetected_chromedriver to avoid detection driver = uc.Chrome(options=chrome_options) # Open the specified website driver.get("https://www.rome2rio.com/map/Marseille/Paris") # Mimicking human behavior with random sleep time.sleep(random.uniform(2, 4)) # Proceed with your script... # Close the driver after operations are complete driver.quit()
优化建议
- 更新并丰富User-Agent池:当前使用的UA都是Chrome 91版本,过于老旧,容易被识别为异常请求。建议添加Chrome 110+、Firefox、Safari等不同浏览器、不同平台(Windows、Mac、Android)的最新UA,每次启动随机选择,模拟真实用户的设备多样性。
- 模拟更贴近人类的交互行为:
- 不要仅依赖固定sleep,结合
ActionChains实现随机鼠标移动(比如从页面左上角随机移动到某个区域)、随机滚动页面(每次滚动100-300像素,间隔0.5-1秒)、偶尔点击空白区域,避免机械性的操作模式。 - 页面加载完成后,先随机等待2-6秒再开始操作,模拟用户浏览页面的习惯。
- 不要仅依赖固定sleep,结合
- 调整undetected_chromedriver配置:
- 移除
--incognito参数,无痕模式的特征反而容易被风控系统标记;改用普通浏览器模式,配合随机用户数据目录(user_data_dir),每次启动使用不同的用户配置文件。 - 随机设置窗口尺寸,替换固定的
--start-maximized,示例代码:chrome_options.add_argument(f"--window-size={random.randint(1200, 1920)},{random.randint(700, 1080)}") - 添加
--disable-extensions、--disable-plugins参数,减少浏览器特征暴露;如果不需要爬取图片,可加入--disable-images降低加载压力和检测概率。
- 移除
- 控制请求频率与间隔:
- 每次请求之间设置5-12秒的随机间隔,避免短时间内发起大量请求。
- 限制每日请求总量,避免触发网站的流量风控机制。
- 使用代理IP轮换:
- 接入高匿代理IP池,每次启动浏览器时切换不同的代理,避免单一IP被封禁。示例代码:
proxies = ["http://proxy1:port", "http://proxy2:port"] chrome_options.add_argument(f'--proxy-server={random.choice(proxies)}')
- 接入高匿代理IP池,每次启动浏览器时切换不同的代理,避免单一IP被封禁。示例代码:
- 备选方案:绕过前端检测:
- 尝试分析网站的API接口,直接发送HTTP请求获取数据,这种方式比模拟浏览器更隐蔽,也更高效。
- 如果仍遇到验证码,可考虑接入验证码识别服务(注意合规性),或手动处理少量验证码请求。
内容的提问来源于stack exchange,提问作者Mateo
相关产品推荐
相关产品推荐

