You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何绕过验证码实现Allegro网页数据爬取?求可行方案

解决Allegro爬取验证码拦截及数据获取方案

一、优化Selenium反检测配置

默认Selenium很容易被平台识别为爬虫,试试这些强化后的配置,能降低验证码触发概率:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.support.wait import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

url = "https://allegro.pl/oferta/lurlego-76900-speed-champions-koenigsegg-jesko-11096977887"

options = webdriver.ChromeOptions()
# 基础运行参数
options.add_argument("--no-sandbox")
options.add_argument("--disable-gpu")
options.add_argument("--window-size=1920x1080")
options.add_argument("--disable-extensions")
# 模拟真实浏览器UA
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
# 禁用自动化标识
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)
options.add_argument("--disable-blink-features=AutomationControlled")

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
# 隐藏webdriver属性,避免被检测
driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")

driver.get(url)
time.sleep(3)

try:
    # 等待标题元素加载完成再获取
    title = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, '//div[@class="msub_k4"]//h4'))
    ).text
    print(title)
except Exception as e:
    print(f"获取失败: {e}")
finally:
    driver.quit()

二、验证码处理办法

如果还是触发验证码,根据你的爬取规模选对应的方案:

  • 手动验证(适合少量爬取):在打开页面后加个等待,手动完成验证码再继续:
    driver.get(url)
    input("完成验证码后按回车继续...")
    title = WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.XPATH, '//div[@class="msub_k4"]//h4'))).text
    
  • 代理IP轮换:频繁请求会触发拦截,用高匿代理池轮换IP,同时控制请求间隔(比如每次请求间隔2-5秒随机时间),减少单IP的请求频率。
  • 第三方识别服务:简单图形验证码可以用Tesseract OCR识别,复杂验证码可以用付费识别平台,但注意这类服务可能违反Allegro用户协议,谨慎使用。

三、最稳定的合规方案:Allegro官方API

直接用平台提供的开发者API,完全不会有验证码问题,而且合法合规:

  1. 去Allegro开发者平台注册账号,创建应用获取API令牌。
  2. 用API直接拉取商品数据,示例代码:
import requests

# 替换成你的API令牌
ACCESS_TOKEN = "your_access_token"
# 从商品URL里提取的商品ID
item_id = "11096977887"

url = f"https://api.allegro.pl/offers/{item_id}"
headers = {
    "Authorization": f"Bearer {ACCESS_TOKEN}",
    "Accept": "application/vnd.allegro.public.v1+json"
}

response = requests.get(url, headers=headers)
if response.status_code == 200:
    item_data = response.json()
    title = item_data["name"]
    print(title)
else:
    print(f"请求失败: {response.status_code}")

这种方式不仅能拿标题,还能获取价格、库存、卖家信息等更多数据,稳定性拉满。

四、额外注意点

  • 定期检查页面元素选择器:Allegro的页面class可能会更新,发现获取失败时先验证XPATH是否有效。
  • 模拟人类行为:不要短时间内批量请求,每次请求加随机间隔,避免被判定为爬虫。

内容的提问来源于stack exchange,提问作者Amen Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 10:01:20