You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取返回0条目0页面,请求技术排查协助

问题排查与修复方案

嘿,看了你的日志和代码,很快发现了两个核心问题,咱们一步步来解决:

首先看日志里的关键提示:

2020-08-29 21:17:30 [scrapy.spidermiddlewares.offsite] DEBUG: Filtered offsite request to 'www.made-in-china.com': <GET https://www.made-in-china.com/multi-search/24v%2Bbattery/F1/2.html;jsessionid=2B77F23449911847145999CD6E9B6429>

这说明Scrapy的**跨站请求中间件(OffsiteMiddleware)**把你的翻页请求直接过滤了,根源是你设置的allowed_domains完全不对。

1. 修正域名白名单

你的代码里写的是:

allowed_domains = ['web']

这显然和目标站made-in-china.com不匹配,改成下面这样就能解决请求被过滤的问题:

allowed_domains = ['made-in-china.com']

2. 重构CrawlSpider的规则配置

现在你的Rule只处理了翻页,还错误地把翻页后的列表页传给了parse_item——但parse_item里的XPath是针对商品详情页的,这就导致:

  • 初始列表页里的商品链接根本没被提取
  • 就算翻页请求没被过滤,也提取不到任何数据

你需要设置两条规则:

  • 第一条:提取列表页中的商品详情链接,调用parse_item解析数据
  • 第二条:提取下一页链接,让Spider自动跟进爬取新的列表页

另外,建议用extract_first()替代extract()来获取单个字段值,避免返回空列表(如果字段不存在的话会返回None,更符合Item的预期)。

修正后的完整Spider代码如下:

import scrapy
from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule
from batt_data.items import BattDataItem

class BatterySpider(CrawlSpider):
    name = 'battery'
    allowed_domains = ['made-in-china.com']  # 修正域名白名单
    start_urls = ['https://www.made-in-china.com/multi-search/24v%2Bbattery/F1/1.html']

    rules = (
        # 规则1:提取商品详情页链接,解析商品数据
        Rule(
            LinkExtractor(restrict_xpaths='//a[@class="product-item-title"]'),  # 你可以在Scrapy Shell里验证这个XPath是否准确
            callback='parse_item',
            follow=False
        ),
        # 规则2:提取下一页链接,继续跟进爬取列表页
        Rule(
            LinkExtractor(restrict_xpaths='//*[contains(@class, "nextpage")]'),
            follow=True
        ),
    )

    def parse_item(self, response):
        item = BattDataItem()
        # 注意:如果XPath在Shell里测试不到数据,需要根据实际页面结构调整
        item['description'] = response.xpath('//img[@class="J-firstLazyload"]/@alt').extract_first()
        item['chemistry'] = response.xpath('//li[@class="J-faketitle ellipsis"][1]/span/text()').extract_first()
        item['applications'] = response.xpath('//li[@class="J-faketitle ellipsis"][2]/span/text()').extract_first()
        item['shape'] = response.xpath('//li[@class="J-faketitle ellipsis"][4]/span/text()').extract_first()
        item['discharge_rate'] = response.xpath('//li[@class="J-faketitle ellipsis"][5]/span/text()').extract_first()
        yield item

额外小贴士

  • 用Scrapy Shell验证XPath:先跑scrapy shell "你的列表页URL",测试商品链接的XPath是否能正确提取;再打开一个详情页,验证字段XPath的有效性。
  • 建议在settings.py里添加DOWNLOAD_DELAY = 2,降低请求频率,避免被网站反爬机制拦截。
  • 检查你的BattDataItem类,确保所有字段都已经正确定义。

内容的提问来源于stack exchange,提问作者Ikponmwosa Endurance

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 23:42:40