You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy同时抓取列表与详情页数据合并及429状态码问题咨询

问题原因与解决方法

1. 列表页与详情页数据分离问题

错误原因

  • 你在parse函数中先后两次独立yield数据:第一次输出列表页信息,第二次发起详情页请求后在parseDetails中单独输出详情页信息,两者没有关联,自然会生成两条独立条目
  • 额外小bug:tmpPrice = tmpPrice.replace("\u00a3",""), 末尾多了个逗号,导致price字段变成了元组类型,对应输出里的["65.00"]格式

解决方法

通过Scrapy的Request.meta属性传递列表页已抓取的信息到详情页回调函数,解析完详情页字段后合并成完整字典再统一输出,不要分开yield。

2. 429状态码报错问题

错误原因

429是服务器限流响应,代表你的请求频率过高触发了站点的反爬机制,不是死循环。你未配置爬取延迟与并发限制,短时间内发起大量请求被站点拦截。

解决方法

修改项目settings.py配置,添加以下规则:

# 设置请求间隔为2-3秒,单位秒
DOWNLOAD_DELAY = 2.5
# 单域名并发请求数调低到2
CONCURRENT_REQUESTS_PER_DOMAIN = 2
# 开启自动限速,会根据响应动态调整请求频率
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 1
AUTOTHROTTLE_MAX_DELAY = 10
# 配置常规浏览器UA,避免默认ScrapyUA被拦截
USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'

修正后完整代码

import scrapy

class WhiskeySpider(scrapy.Spider):
  name = "whiskyDetail"
  allowed_domains = ["whiskyshop.com"]
  start_urls = ["https://www.whiskyshop.com/scotch-whisky"]

  def parse(self, response):
    for products in response.css("div.product-item-info"):
      tmpPrice = products.css("span.price::text").get()      
      tmpLink = products.css("a.product-item-link").attrib["href"]
      tmpLink = response.urljoin(tmpLink)
      
      # 先组装列表页已获取的信息
      item = {
        "name": products.css("a.product-item-link::text").get().strip(),
        "price": tmpPrice.replace("\u00a3","") if tmpPrice else "Sold Out", # 去掉多余逗号,修复元组问题
        "link": tmpLink,
      }
      # 把列表页数据通过meta传给详情页回调,不再单独yield
      yield scrapy.Request(url=tmpLink, callback=self.parseDetails, meta={"item": item})                    
    
    nextPage = response.css("a.action.next").attrib.get("href") # 用get避免无下一页时报KeyError
    if nextPage:
      nextPage = response.urljoin(nextPage)
      yield response.follow(nextPage, callback=self.parse)
  
  def parseDetails(self, response):
    # 取出meta传递的列表页数据
    item = response.meta["item"]
    tmpDetails = response.css("p.product-info-size-abv span::text").getall()
    # 补充详情页字段,做兼容处理避免索引报错
    if len(tmpDetails) >=3:
      item["litre"] = tmpDetails[0].strip()
      item["percent"] = tmpDetails[1].strip()
      item["area"] = tmpDetails[2].strip()
    else:
      item["litre"] = item["percent"] = item["area"] = "未获取到"
    # 统一yield合并后的完整数据
    yield item

内容的提问来源于stack exchange,提问作者Rapid1898

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 04:15:02