Scrapy爬取amazon.de时修改配送地址位置不生效问题求助
Scrapy爬取amazon.de地址修改失效排查方案
核心问题排查点
- 缺失CSRF校验参数
亚马逊的/gp/delivery/ajax/address-change.html接口做了CSRF防护,你当前直接提交表单未携带校验用的token,请求本质上不会被后端受理。需要先从首次请求亚马逊首页的响应里提取两类关键参数:- 响应头或者页面隐藏标签中的
anti-csrftoken-a2z值,需要添加到地址修改请求的请求头中 - 部分场景下需要携带表单参数
csrfmiddlewaretoken,值同样从首页源码中提取
- 响应头或者页面隐藏标签中的
- Cookie持久化配置错误
要确认Scrapy全局配置settings.py中COOKIES_ENABLED = True,只有开启cookie持久化,地址修改接口返回的地址标记cookie(session-id、ubid-acbde、i18n-prefs等)才会被后续请求自动携带,否则地址修改的状态不会被保留。 - 请求头伪装不足
亚马逊会对缺省Scrapy请求头的请求直接返回兜底的通用数据,不会生效地址修改配置,地址修改请求需要补充Accept、X-Requested-With、Referer等标准浏览器请求头,同时替换User-Agent为真实浏览器的UA字符串。 - 表单参数不全
当前提交的表单可以补充countryCode: DE参数,明确指定邮编所属国家,避免接口识别邮编归属错误。
修正后核心代码示例
import scrapy from scrapy.spiders import Spider as BaseSpider class AmazonSpider(BaseSpider): name = 'amazon' allowed_domains = ['www.amazon.de'] start_urls = ['https://www.amazon.de/'] custom_settings = { 'COOKIES_ENABLED': True, 'DEFAULT_REQUEST_HEADERS': { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'de-DE,de;q=0.9,en-US;q=0.8,en;q=0.7', 'Referer': 'https://www.amazon.de/' } } def parse(self, response): # 提取CSRF校验参数 csrf_token = response.xpath('//input[@name="csrfmiddlewaretoken"]/@value').get() anti_csrf = response.headers.get('anti-csrftoken-a2z', b'').decode() if not anti_csrf: anti_csrf = response.xpath('//meta[@name="csrf-token"]/@content').get('') data = { 'locationType': 'LOCATION_INPUT', 'zipCode': '10115', 'countryCode': 'DE', 'storeContext': 'drugstore', 'deviceType': 'web', 'pageType': 'Detail', 'actionSource': 'glow', 'almBrandId': 'undefined', 'csrfmiddlewaretoken': csrf_token } headers = { 'Accept': 'application/json, text/javascript, */*; q=0.01', 'X-Requested-With': 'XMLHttpRequest', 'anti-csrftoken-a2z': anti_csrf } yield scrapy.FormRequest( url='https://www.amazon.de/gp/delivery/ajax/address-change.html', formdata=data, headers=headers, callback=self.parse_pages ) def parse_pages(self, response): # 校验地址修改是否生效 if 'isValidAddress":true' in response.text: url = 'https://www.amazon.de/-/en/Filter-Computer-Glasses-Headache-Vintage/dp/B091FYYDXB/ref=sr_1_95?dchild=1&keywords=kopfschmerzen&qid=1630410090&s=drugstore&sr=1-95' yield response.follow( url=url, dont_filter=True, callback=self.parse_product ) def parse_product(self, response): # 商品解析逻辑 pass
内容的提问来源于stack exchange,提问作者Roman
相关产品推荐
相关产品推荐

