Scrapy爬虫运行但未抓取页面,山地车网站爬取无数据求助
我是Web Scraping新手,尝试编写单脚本Scrapy Spider,从山地车销售网站抓取商品名称、品牌及价格信息。爬虫可正常运行,但生成的bikes.csv文件为空,终端显示日志:INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min),无法确定未抓取页面及数据的原因。
我已测试代码中的URL、定位目标信息的CSS选择器(也尝试过Xpath但无效),还修改过爬取全页的循环逻辑,但均未解决问题。目前怀疑代码开头存在语法错误导致爬虫异常,或翻页循环逻辑有问题,另外该网站采用无限滚动模式,这是否会影响爬取?
相关代码如下:
import scrapy import requests from scrapy.crawler import CrawlerProcess class BikeSpider(scrapy.Spider): name='mountianbikespider' def start_requests(self): yield scrapy.Request('https://www.incycle.com/pages/search-results-page?collection=mountain-bikes&page=1') def parse(self, response): products = response.css('li.snize-product-in-stock') for item in products: yield { 'name' : item.css('span.snistrong textze-title::text').extract(), 'description' : item.css('span.snize-description::text').extract(), 'price' : item.css('span.snize-price::text').extract() } #this loop will make the spider not only crawl the first page of bikes, but also continue to all pages afterwards, collection the same info as on page 1 # to do this you must change the url to include page={x} in the place of page=1 for x in range(2,10): yield(scrapy.Request(f'https://www.incycle.com/pages/search-results-page?collection=mountain-bikes&page={x}', callback=self.parse)) #this is what saves the data in a seperate place (in this case a csv namesbikes.csv) process = CrawlerProcess(settings={ "FEEDS":{"bikes.csv":{"format": "csv"}} }) #this is what actually runs the spider process.crawl(BikeSpider) process.start()
1. 修正CSS选择器拼写错误
你的商品名称选择器存在明显笔误:
- 错误写法:
span.snistrong textze-title::text - 正确写法:
span.snize-strong.snize-title::text(页面实际类名为snize-strong和snize-title,并非你写的错误拼写)
同时,建议用get().strip()替代extract(),避免返回空列表,修改后的item提取代码:
yield { 'name': item.css('span.snize-strong.snize-title::text').get(default='').strip(), 'description': item.css('span.snize-description::text').get(default='').strip(), 'price': item.css('span.snize-price::text').get(default='').strip() }
2. 检查并修正缩进逻辑
你的翻页循环缩进存在问题,当前代码中它的缩进层级与商品循环内部的yield对齐,导致翻页逻辑无法执行。需要将翻页循环调整到与products = response.css(...)同一缩进层级,确保在商品提取完成后触发翻页请求。
3. 处理动态渲染与反爬拦截
该网站采用无限滚动,说明数据大概率是动态加载的;同时Scrapy默认请求头可能被网站识别为爬虫,导致返回空页面。可以通过以下方式解决:
- 添加浏览器模拟请求头:
在start_requests和翻页请求中加入自定义User-Agent:headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } - 如果页面完全由JavaScript渲染,Scrapy默认无法获取内容,需要使用
scrapy-playwright或selenium处理动态页面。
4. 验证请求有效性
手动访问你构造的带page参数的URL,查看页面源码是否包含li.snize-product-in-stock元素。如果源码中没有对应内容,说明该URL并非真实数据接口,需要通过浏览器抓包找到实际的AJAX接口,直接爬取接口返回的JSON数据。
5. 清理冗余代码
删除未使用的import requests语句,减少不必要的资源加载。
import scrapy from scrapy.crawler import CrawlerProcess class BikeSpider(scrapy.Spider): name='mountianbikespider' def start_requests(self): headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } yield scrapy.Request( 'https://www.incycle.com/pages/search-results-page?collection=mountain-bikes&page=1', headers=headers ) def parse(self, response): products = response.css('li.snize-product-in-stock') # 日志输出找到的商品数量,方便排查 self.logger.info(f'当前页面找到 {len(products)} 个商品') for item in products: yield { 'name': item.css('span.snize-strong.snize-title::text').get(default='').strip(), 'description': item.css('span.snize-description::text').get(default='').strip(), 'price': item.css('span.snize-price::text').get(default='').strip() } # 翻页逻辑,调整到正确缩进层级 for x in range(2,10): headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } yield scrapy.Request( f'https://www.incycle.com/pages/search-results-page?collection=mountain-bikes&page={x}', headers=headers, callback=self.parse ) process = CrawlerProcess(settings={ "FEEDS": {"bikes.csv": {"format": "csv"}} }) process.crawl(BikeSpider) process.start()
内容的提问来源于stack exchange,提问作者Anatole Colevas

