You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫运行但未抓取页面,山地车网站爬取无数据求助

问题描述

我是Web Scraping新手,尝试编写单脚本Scrapy Spider,从山地车销售网站抓取商品名称、品牌及价格信息。爬虫可正常运行,但生成的bikes.csv文件为空,终端显示日志:INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min),无法确定未抓取页面及数据的原因。

我已测试代码中的URL、定位目标信息的CSS选择器(也尝试过Xpath但无效),还修改过爬取全页的循环逻辑,但均未解决问题。目前怀疑代码开头存在语法错误导致爬虫异常,或翻页循环逻辑有问题,另外该网站采用无限滚动模式,这是否会影响爬取?

相关代码如下:

import scrapy
import requests
from scrapy.crawler import CrawlerProcess

class BikeSpider(scrapy.Spider):
    name='mountianbikespider'
    
    def start_requests(self):
        yield scrapy.Request('https://www.incycle.com/pages/search-results-page?collection=mountain-bikes&page=1')
        
    def parse(self, response):
        products = response.css('li.snize-product-in-stock')
        for item in products:      
            yield {
                'name' : item.css('span.snistrong textze-title::text').extract(),
                'description' : item.css('span.snize-description::text').extract(),
                'price' : item.css('span.snize-price::text').extract()    
            }
            
#this loop will make the spider not only crawl the first page of bikes, but also continue to all pages afterwards, collection the same info as on page 1
# to do this you must change the url to include page={x} in the place of page=1
        for x in range(2,10):
            yield(scrapy.Request(f'https://www.incycle.com/pages/search-results-page?collection=mountain-bikes&page={x}', callback=self.parse))
        
#this is what saves the data in a seperate place (in this case a csv namesbikes.csv)
process = CrawlerProcess(settings={
    "FEEDS":{"bikes.csv":{"format": "csv"}} 
})    

#this is what actually runs the spider
process.crawl(BikeSpider)
process.start()

问题分析与解决办法

1. 修正CSS选择器拼写错误

你的商品名称选择器存在明显笔误:

  • 错误写法:span.snistrong textze-title::text
  • 正确写法:span.snize-strong.snize-title::text(页面实际类名为snize-strong和snize-title,并非你写的错误拼写)

同时,建议用get().strip()替代extract(),避免返回空列表,修改后的item提取代码:

yield {
    'name': item.css('span.snize-strong.snize-title::text').get(default='').strip(),
    'description': item.css('span.snize-description::text').get(default='').strip(),
    'price': item.css('span.snize-price::text').get(default='').strip()
}

2. 检查并修正缩进逻辑

你的翻页循环缩进存在问题,当前代码中它的缩进层级与商品循环内部的yield对齐,导致翻页逻辑无法执行。需要将翻页循环调整到与products = response.css(...)同一缩进层级,确保在商品提取完成后触发翻页请求。

3. 处理动态渲染与反爬拦截

该网站采用无限滚动,说明数据大概率是动态加载的;同时Scrapy默认请求头可能被网站识别为爬虫,导致返回空页面。可以通过以下方式解决:

  • 添加浏览器模拟请求头:
    在start_requests和翻页请求中加入自定义User-Agent:
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    
  • 如果页面完全由JavaScript渲染,Scrapy默认无法获取内容,需要使用scrapy-playwright或selenium处理动态页面。

4. 验证请求有效性

手动访问你构造的带page参数的URL,查看页面源码是否包含li.snize-product-in-stock元素。如果源码中没有对应内容,说明该URL并非真实数据接口,需要通过浏览器抓包找到实际的AJAX接口,直接爬取接口返回的JSON数据。

5. 清理冗余代码

删除未使用的import requests语句,减少不必要的资源加载。


修改后的完整代码示例
import scrapy
from scrapy.crawler import CrawlerProcess

class BikeSpider(scrapy.Spider):
    name='mountianbikespider'
    
    def start_requests(self):
        headers = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
        }
        yield scrapy.Request(
            'https://www.incycle.com/pages/search-results-page?collection=mountain-bikes&page=1',
            headers=headers
        )
        
    def parse(self, response):
        products = response.css('li.snize-product-in-stock')
        # 日志输出找到的商品数量,方便排查
        self.logger.info(f'当前页面找到 {len(products)} 个商品')
        for item in products:      
            yield {
                'name': item.css('span.snize-strong.snize-title::text').get(default='').strip(),
                'description': item.css('span.snize-description::text').get(default='').strip(),
                'price': item.css('span.snize-price::text').get(default='').strip()
            }
            
        # 翻页逻辑,调整到正确缩进层级
        for x in range(2,10):
            headers = {
                'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
            }
            yield scrapy.Request(
                f'https://www.incycle.com/pages/search-results-page?collection=mountain-bikes&page={x}',
                headers=headers,
                callback=self.parse
            )
        
process = CrawlerProcess(settings={
    "FEEDS": {"bikes.csv": {"format": "csv"}} 
})    

process.crawl(BikeSpider)
process.start()

内容的提问来源于stack exchange,提问作者Anatole Colevas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 06:55:21