Scrapy Spider跳转URL失败,求助获取标准棚屋产品价格
Scrapy爬虫爬取产品价格问题解决
问题梳理
需要爬取standard-sheds分类页(https://www.charnleys.co.uk/product-category/gardening/garden-accessories/garden-furniture/sheds/standard-sheds/)下所有产品的价格,但产品详情页路径为/shop/[产品名],现有代码无法追踪跳转,且不清楚如何遍历产品URL数组实现详情页爬取。
现有代码的问题
- 全局
urls数组不适合Scrapy异步架构,容易出现数据混乱 - Rules存在缩进错误,规则逻辑不合理,未正确引导爬虫到达目标分类页
collect_urls仅收集URL但未生成详情页爬取请求html_return_price_strings仅打印价格,无返回值,解析方式低效parse_product参数错误,Scrapy回调函数无法直接传递自定义方法作为参数
修正后的代码(CrawlSpider版本)
import scrapy from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor class CharnleySpider(CrawlSpider): name = 'crawler' allowed_domains = ['charnleys.co.uk'] # 直接从目标分类页开始爬取,减少无效遍历 start_urls = ['https://www.charnleys.co.uk/product-category/gardening/garden-accessories/garden-furniture/sheds/standard-sheds/'] rules = ( # 匹配所有/shop/开头的详情页链接,直接回调解析方法 Rule(LinkExtractor(allow=r'/shop/'), callback='parse_product', follow=False), # 处理分类页分页(如果存在分页则自动跟进) Rule(LinkExtractor(allow=r'page/\d+'), follow=True), ) def parse_product(self, response): # 直接通过CSS选择器提取产品名称和价格 yield { 'name': response.css('h2.product_title::text').get().strip(), 'price': response.css('p.price span.woocommerce-Price-amount::text').get().strip() }
修正后的代码(普通Spider版本)
如果更倾向手动控制爬取流程,可使用普通Spider实现:
import scrapy class CharnleySpider(scrapy.Spider): name = 'crawler' allowed_domains = ['charnleys.co.uk'] start_urls = ['https://www.charnleys.co.uk/product-category/gardening/garden-accessories/garden-furniture/sheds/standard-sheds/'] def parse(self, response): # 提取分类页所有产品链接 product_links = response.css('div.product-image a::attr(href)').getall() # 遍历链接生成详情页请求 for link in product_links: yield scrapy.Request(url=link, callback=self.parse_product) # 处理分页(如果有下一页则继续爬取) next_page = response.css('a.next::attr(href)').get() if next_page: yield scrapy.Request(url=next_page, callback=self.parse) def parse_product(self, response): yield { 'name': response.css('h2.product_title::text').get().strip(), 'price': response.css('p.price span.woocommerce-Price-amount::text').get().strip() }
关键要点
- Scrapy的异步架构不需要手动维护URL数组,通过
yield scrapy.Request或CrawlSpider的Rules即可自动管理请求队列 - 优先使用CSS/XPath选择器直接定位目标元素,避免遍历整个HTML,大幅提升解析效率
- 如果分类页存在分页,必须添加分页处理逻辑,否则只能爬取第一页的产品
内容的提问来源于stack exchange,提问作者Bastion
相关产品推荐
相关产品推荐

