Scrapy爬取返回0条目0页面,请求技术排查协助
问题排查与修复方案
嘿,看了你的日志和代码,很快发现了两个核心问题,咱们一步步来解决:
首先看日志里的关键提示:
2020-08-29 21:17:30 [scrapy.spidermiddlewares.offsite] DEBUG: Filtered offsite request to 'www.made-in-china.com': <GET https://www.made-in-china.com/multi-search/24v%2Bbattery/F1/2.html;jsessionid=2B77F23449911847145999CD6E9B6429>
这说明Scrapy的**跨站请求中间件(OffsiteMiddleware)**把你的翻页请求直接过滤了,根源是你设置的allowed_domains完全不对。
1. 修正域名白名单
你的代码里写的是:
allowed_domains = ['web']
这显然和目标站made-in-china.com不匹配,改成下面这样就能解决请求被过滤的问题:
allowed_domains = ['made-in-china.com']
2. 重构CrawlSpider的规则配置
现在你的Rule只处理了翻页,还错误地把翻页后的列表页传给了parse_item——但parse_item里的XPath是针对商品详情页的,这就导致:
- 初始列表页里的商品链接根本没被提取
- 就算翻页请求没被过滤,也提取不到任何数据
你需要设置两条规则:
- 第一条:提取列表页中的商品详情链接,调用
parse_item解析数据 - 第二条:提取下一页链接,让Spider自动跟进爬取新的列表页
另外,建议用extract_first()替代extract()来获取单个字段值,避免返回空列表(如果字段不存在的话会返回None,更符合Item的预期)。
修正后的完整Spider代码如下:
import scrapy from scrapy.linkextractors import LinkExtractor from scrapy.spiders import CrawlSpider, Rule from batt_data.items import BattDataItem class BatterySpider(CrawlSpider): name = 'battery' allowed_domains = ['made-in-china.com'] # 修正域名白名单 start_urls = ['https://www.made-in-china.com/multi-search/24v%2Bbattery/F1/1.html'] rules = ( # 规则1:提取商品详情页链接,解析商品数据 Rule( LinkExtractor(restrict_xpaths='//a[@class="product-item-title"]'), # 你可以在Scrapy Shell里验证这个XPath是否准确 callback='parse_item', follow=False ), # 规则2:提取下一页链接,继续跟进爬取列表页 Rule( LinkExtractor(restrict_xpaths='//*[contains(@class, "nextpage")]'), follow=True ), ) def parse_item(self, response): item = BattDataItem() # 注意:如果XPath在Shell里测试不到数据,需要根据实际页面结构调整 item['description'] = response.xpath('//img[@class="J-firstLazyload"]/@alt').extract_first() item['chemistry'] = response.xpath('//li[@class="J-faketitle ellipsis"][1]/span/text()').extract_first() item['applications'] = response.xpath('//li[@class="J-faketitle ellipsis"][2]/span/text()').extract_first() item['shape'] = response.xpath('//li[@class="J-faketitle ellipsis"][4]/span/text()').extract_first() item['discharge_rate'] = response.xpath('//li[@class="J-faketitle ellipsis"][5]/span/text()').extract_first() yield item
额外小贴士
- 用Scrapy Shell验证XPath:先跑
scrapy shell "你的列表页URL",测试商品链接的XPath是否能正确提取;再打开一个详情页,验证字段XPath的有效性。 - 建议在
settings.py里添加DOWNLOAD_DELAY = 2,降低请求频率,避免被网站反爬机制拦截。 - 检查你的
BattDataItem类,确保所有字段都已经正确定义。
内容的提问来源于stack exchange,提问作者Ikponmwosa Endurance
相关产品推荐
相关产品推荐

