如何绕过验证码实现Allegro网页数据爬取?求可行方案
解决Allegro爬取验证码拦截及数据获取方案
一、优化Selenium反检测配置
默认Selenium很容易被平台识别为爬虫,试试这些强化后的配置,能降低验证码触发概率:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.support.wait import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time url = "https://allegro.pl/oferta/lurlego-76900-speed-champions-koenigsegg-jesko-11096977887" options = webdriver.ChromeOptions() # 基础运行参数 options.add_argument("--no-sandbox") options.add_argument("--disable-gpu") options.add_argument("--window-size=1920x1080") options.add_argument("--disable-extensions") # 模拟真实浏览器UA options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") # 禁用自动化标识 options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) options.add_argument("--disable-blink-features=AutomationControlled") driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) # 隐藏webdriver属性,避免被检测 driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})") driver.get(url) time.sleep(3) try: # 等待标题元素加载完成再获取 title = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, '//div[@class="msub_k4"]//h4')) ).text print(title) except Exception as e: print(f"获取失败: {e}") finally: driver.quit()
二、验证码处理办法
如果还是触发验证码,根据你的爬取规模选对应的方案:
- 手动验证(适合少量爬取):在打开页面后加个等待,手动完成验证码再继续:
driver.get(url) input("完成验证码后按回车继续...") title = WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.XPATH, '//div[@class="msub_k4"]//h4'))).text - 代理IP轮换:频繁请求会触发拦截,用高匿代理池轮换IP,同时控制请求间隔(比如每次请求间隔2-5秒随机时间),减少单IP的请求频率。
- 第三方识别服务:简单图形验证码可以用Tesseract OCR识别,复杂验证码可以用付费识别平台,但注意这类服务可能违反Allegro用户协议,谨慎使用。
三、最稳定的合规方案:Allegro官方API
直接用平台提供的开发者API,完全不会有验证码问题,而且合法合规:
- 去Allegro开发者平台注册账号,创建应用获取API令牌。
- 用API直接拉取商品数据,示例代码:
import requests # 替换成你的API令牌 ACCESS_TOKEN = "your_access_token" # 从商品URL里提取的商品ID item_id = "11096977887" url = f"https://api.allegro.pl/offers/{item_id}" headers = { "Authorization": f"Bearer {ACCESS_TOKEN}", "Accept": "application/vnd.allegro.public.v1+json" } response = requests.get(url, headers=headers) if response.status_code == 200: item_data = response.json() title = item_data["name"] print(title) else: print(f"请求失败: {response.status_code}")
这种方式不仅能拿标题,还能获取价格、库存、卖家信息等更多数据,稳定性拉满。
四、额外注意点
- 定期检查页面元素选择器:Allegro的页面class可能会更新,发现获取失败时先验证XPATH是否有效。
- 模拟人类行为:不要短时间内批量请求,每次请求加随机间隔,避免被判定为爬虫。
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

