Scrapy同时抓取列表与详情页数据合并及429状态码问题咨询
问题原因与解决方法
1. 列表页与详情页数据分离问题
错误原因
- 你在
parse函数中先后两次独立yield数据:第一次输出列表页信息,第二次发起详情页请求后在parseDetails中单独输出详情页信息,两者没有关联,自然会生成两条独立条目 - 额外小bug:
tmpPrice = tmpPrice.replace("\u00a3",""),末尾多了个逗号,导致price字段变成了元组类型,对应输出里的["65.00"]格式
解决方法
通过Scrapy的Request.meta属性传递列表页已抓取的信息到详情页回调函数,解析完详情页字段后合并成完整字典再统一输出,不要分开yield。
2. 429状态码报错问题
错误原因
429是服务器限流响应,代表你的请求频率过高触发了站点的反爬机制,不是死循环。你未配置爬取延迟与并发限制,短时间内发起大量请求被站点拦截。
解决方法
修改项目settings.py配置,添加以下规则:
# 设置请求间隔为2-3秒,单位秒 DOWNLOAD_DELAY = 2.5 # 单域名并发请求数调低到2 CONCURRENT_REQUESTS_PER_DOMAIN = 2 # 开启自动限速,会根据响应动态调整请求频率 AUTOTHROTTLE_ENABLED = True AUTOTHROTTLE_START_DELAY = 1 AUTOTHROTTLE_MAX_DELAY = 10 # 配置常规浏览器UA,避免默认ScrapyUA被拦截 USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
修正后完整代码
import scrapy class WhiskeySpider(scrapy.Spider): name = "whiskyDetail" allowed_domains = ["whiskyshop.com"] start_urls = ["https://www.whiskyshop.com/scotch-whisky"] def parse(self, response): for products in response.css("div.product-item-info"): tmpPrice = products.css("span.price::text").get() tmpLink = products.css("a.product-item-link").attrib["href"] tmpLink = response.urljoin(tmpLink) # 先组装列表页已获取的信息 item = { "name": products.css("a.product-item-link::text").get().strip(), "price": tmpPrice.replace("\u00a3","") if tmpPrice else "Sold Out", # 去掉多余逗号,修复元组问题 "link": tmpLink, } # 把列表页数据通过meta传给详情页回调,不再单独yield yield scrapy.Request(url=tmpLink, callback=self.parseDetails, meta={"item": item}) nextPage = response.css("a.action.next").attrib.get("href") # 用get避免无下一页时报KeyError if nextPage: nextPage = response.urljoin(nextPage) yield response.follow(nextPage, callback=self.parse) def parseDetails(self, response): # 取出meta传递的列表页数据 item = response.meta["item"] tmpDetails = response.css("p.product-info-size-abv span::text").getall() # 补充详情页字段,做兼容处理避免索引报错 if len(tmpDetails) >=3: item["litre"] = tmpDetails[0].strip() item["percent"] = tmpDetails[1].strip() item["area"] = tmpDetails[2].strip() else: item["litre"] = item["percent"] = item["area"] = "未获取到" # 统一yield合并后的完整数据 yield item
内容的提问来源于stack exchange,提问作者Rapid1898
相关产品推荐
相关产品推荐

