使用Requests爬取亚马逊商品价格遇反爬错误,求解决方法
我尝试使用Requests库爬取亚马逊商品页面提取价格,但始终收到以下反爬提示:
"如需讨论亚马逊数据的自动化访问事宜,请联系api-services-support@amazon.com。
有关迁移至我们API的信息,请参考我们的Marketplace APIs,或针对广告场景的Product Advertising API。
抱歉,我们需要确认您并非机器人。为获得最佳效果,请确保您的浏览器接受Cookie。"
以下是我当前使用的代码:import requests header = { "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7", "Accept-Encoding": "gzip, deflate, br", "Accept-Language": "en-US,en;q=0.9", "Upgrade-Insecure-Requests": "1", "Referer": "https://www.google.com/", "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } url = 'https://www.amazon.com/-/es/Revlon-One-Step-Volumizer-PLUS/dp/B096SVJZSW/?_encoding=UTF8&content-id=amzn1.sym.3f4ca281-e55c-46d1-9425-fb252d20366f&ref_=pd_gw_exports_top_sellers_unrec' response = requests.get(url, headers=header) data=response.text print(data) print(response.status_code)
解决办法
1. 维持Cookie会话
亚马逊依赖Cookie验证用户身份,单条requests.get不会保留会话状态,改用requests.Session()持久化Cookie:
session = requests.Session() response = session.get(url, headers=header)
2. 补充完整请求头
当前请求头缺少亚马逊识别的关键字段,从真实浏览器的网络请求中复制完整请求头替换现有内容,至少补充Connection和Cache-Control:
header = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7", "Accept-Encoding": "gzip, deflate, br", "Accept-Language": "en-US,en;q=0.9", "Upgrade-Insecure-Requests": "1", "Referer": "https://www.amazon.com/", "Connection": "keep-alive", "Cache-Control": "max-age=0" }
3. 控制请求频率
频繁请求会触发反爬机制,每次请求后添加1-3秒的随机延迟:
import time import random # 请求后随机延迟 time.sleep(random.uniform(1, 3))
4. 合规方案:使用官方API
如果是商业用途,直接使用亚马逊提供的Marketplace APIs或Product Advertising API,这是合法且稳定的数据获取方式,完全不会触发反爬。
5. 模拟真实浏览器(备选)
如果以上方法无效,用selenium或playwright模拟浏览器行为,绕过基础反爬检测:
from selenium import webdriver from selenium.webdriver.chrome.options import Options options = Options() options.add_argument("--headless=new") options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=options) driver.get(url) print(driver.page_source) driver.quit()
内容的提问来源于stack exchange,提问作者chunko
相关产品推荐
相关产品推荐

