亚马逊商品解析问题:503错误及解析结果不稳定求助
问题背景
你遇到的503 Service Unavailable错误以及解析结果时好时坏、字段为Null的情况,完全是亚马逊反爬机制导致的——要么请求被拦截,要么返回的页面是反爬页面(比如空白页、验证码页),导致解析失败。
错误日志
Traceback (most recent call last): File "c:\Users\grafity\Desktop\main.py", line 42, in main() File "c:\Users\grafity\Desktop\main.py", line 31, in main products = get_products(search_url) ^^^^^^^^^^^^^^^^^^^^^^^^ File "c:\Users\grafity\Desktop\main.py", line 14, in get_products response.raise_for_status() File "C:\Users\grafity\AppData\Local\Programs\Python\Python311\Lib\site-packages\requests\models.py", line 1021, in raise_for_status raise HTTPError(http_error_msg, response=self) requests.exceptions.HTTPError: 503 Server Error: Service Unavailable for url: https://www.amazon.ca/s?k=RTX+3070ti
不稳定解析示例
{"img": null, "name": null, "price": null, "url": null}, {"img": "https://m.media-amazon.com/images/I/71vWKkde2yL._AC_UY218_.jpg", "name": null, "price": null, "url": "/sspa/click?ie=UTF8&spc=MToyNjUzMzM1MzIwODAzNjE5OjE2OTQ1MzczOTE6c3BfYXRmOjMwMDAwMTUzMzIyNDgwMjo6MDo6&url=%2FPNY-GeForce-Uprising-Triple-Graphics%2Fdp%2FB0C1HVL3BD%2Fref%3Dsr_1_1_sspa%3Fkeywords%3DRTX%2B3070ti%26qid%3D1694537391%26sr%3D8-1-spons%26sp_csd%3Dd2lkZ2V0TmFtZT1zcF9hdGY%26psc%3D1"}, {"img": "https://m.media-amazon.com/images/I/81d0pk3xOiS._AC_UY218_.jpg", "name": null, "price": 843.0, "url": "/MSI-Gaming-RTX-3070-Trio/dp/B096HJC18P/ref=sr_1_2?keywords=RTX+3070ti&qid=1694537391&sr=8-2"}
具体解决措施
模拟真实浏览器请求头
不要用requests默认的请求头,直接复制Chrome/Firefox浏览器的真实请求头,至少包含User-Agent、Accept、Accept-Language字段。可以维护一个User-Agent列表,每次请求随机选一个,避免固定标识被识别。示例:import random user_agents = [ "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36", "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_5) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.5 Safari/605.1.15" ] headers = { "User-Agent": random.choice(user_agents), "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8", "Accept-Language": "en-CA,en-US;q=0.7,en;q=0.3" } response = requests.get(url, headers=headers)严格控制请求频率
亚马逊对请求频率敏感,每次请求后加随机延迟(比如2-5秒),不要短时间内连续发起请求。避免用循环批量请求,改成异步请求也要控制并发数(最多2-3个并发)。示例:import time time.sleep(random.uniform(2, 5))使用代理IP规避封锁
如果本地IP被封,换用代理IP,优先选住宅代理(Residential Proxy),这类IP更接近真实用户,不容易被识别。每次请求随机切换代理,避免单一IP被频繁检测。改用浏览器渲染工具处理动态内容
亚马逊很多商品数据是通过JavaScript动态加载的,纯requests静态抓取可能拿不到完整页面。用Selenium或Playwright模拟真实浏览器加载页面,能获取到渲染后的完整HTML,减少解析时的Null字段。示例(Playwright):from playwright.sync_api import sync_playwright with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page() page.set_extra_http_headers({"User-Agent": random.choice(user_agents)}) page.goto(search_url, wait_until="networkidle") html = page.content() # 后续解析html browser.close()完善错误处理与重试机制
捕获503、403等反爬错误,设置重试逻辑(最多3-5次重试),重试前加更长延迟。解析前先判断页面是否正常(比如检查是否包含亚马逊商品卡片的特征元素),如果是反爬页面(比如有验证码提示),直接跳过或重试。示例:from requests.exceptions import HTTPError def get_products(url): max_retries = 3 for attempt in range(max_retries): try: response = requests.get(url, headers=headers) response.raise_for_status() # 检查页面是否正常,比如找商品卡片的父容器 if "s-result-list" not in response.text: raise ValueError("Anti-scraping page returned") return parse_html(response.text) except (HTTPError, ValueError) as e: if attempt == max_retries -1: raise time.sleep(random.uniform(5, 10))优化解析逻辑的健壮性
不要依赖固定的HTML类名或结构(亚马逊经常修改页面结构),用更通用的选择器:- 先定位所有商品卡片的共同父元素(比如
div[data-component-type="s-search-result"]) - 对每个商品卡片,分别查找img、name、price元素,用判断逻辑避免元素不存在时报错
- 价格解析时,处理多种格式(比如
$843.00、CAD 843),提取数字部分
示例(用BeautifulSoup):
from bs4 import BeautifulSoup def parse_html(html): soup = BeautifulSoup(html, "html.parser") products = [] for item in soup.select('div[data-component-type="s-search-result"]'): product = { "img": item.select_one("img.s-image")["src"] if item.select_one("img.s-image") else None, "name": item.select_one("h2 a span").get_text(strip=True) if item.select_one("h2 a span") else None, "price": float(item.select_one(".a-price-whole").get_text(strip=True)) if item.select_one(".a-price-whole") else None, "url": "https://www.amazon.ca" + item.select_one("h2 a")["href"] if item.select_one("h2 a") else None } products.append(product) return products- 先定位所有商品卡片的共同父元素(比如
避免触发验证码
尽量不要在短时间内用同一个IP搜索相同关键词,不要频繁切换关键词或地区。如果遇到验证码,可以手动处理后继续,或者暂时停止请求一段时间。
内容的提问来源于stack exchange,提问作者Timur

