You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy Spider跳转URL失败,求助获取标准棚屋产品价格

Scrapy爬虫爬取产品价格问题解决

问题梳理

需要爬取standard-sheds分类页(https://www.charnleys.co.uk/product-category/gardening/garden-accessories/garden-furniture/sheds/standard-sheds/)下所有产品的价格,但产品详情页路径为/shop/[产品名],现有代码无法追踪跳转,且不清楚如何遍历产品URL数组实现详情页爬取。

现有代码的问题

  • 全局urls数组不适合Scrapy异步架构,容易出现数据混乱
  • Rules存在缩进错误,规则逻辑不合理,未正确引导爬虫到达目标分类页
  • collect_urls仅收集URL但未生成详情页爬取请求
  • html_return_price_strings仅打印价格,无返回值,解析方式低效
  • parse_product参数错误,Scrapy回调函数无法直接传递自定义方法作为参数

修正后的代码(CrawlSpider版本)

import scrapy
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor

class CharnleySpider(CrawlSpider):
    name = 'crawler'
    allowed_domains = ['charnleys.co.uk']
    # 直接从目标分类页开始爬取,减少无效遍历
    start_urls = ['https://www.charnleys.co.uk/product-category/gardening/garden-accessories/garden-furniture/sheds/standard-sheds/']

    rules = (
        # 匹配所有/shop/开头的详情页链接,直接回调解析方法
        Rule(LinkExtractor(allow=r'/shop/'), callback='parse_product', follow=False),
        # 处理分类页分页(如果存在分页则自动跟进)
        Rule(LinkExtractor(allow=r'page/\d+'), follow=True),
    )

    def parse_product(self, response):
        # 直接通过CSS选择器提取产品名称和价格
        yield {
            'name': response.css('h2.product_title::text').get().strip(),
            'price': response.css('p.price span.woocommerce-Price-amount::text').get().strip()
        }

修正后的代码(普通Spider版本)

如果更倾向手动控制爬取流程,可使用普通Spider实现:

import scrapy

class CharnleySpider(scrapy.Spider):
    name = 'crawler'
    allowed_domains = ['charnleys.co.uk']
    start_urls = ['https://www.charnleys.co.uk/product-category/gardening/garden-accessories/garden-furniture/sheds/standard-sheds/']

    def parse(self, response):
        # 提取分类页所有产品链接
        product_links = response.css('div.product-image a::attr(href)').getall()
        # 遍历链接生成详情页请求
        for link in product_links:
            yield scrapy.Request(url=link, callback=self.parse_product)
        
        # 处理分页(如果有下一页则继续爬取)
        next_page = response.css('a.next::attr(href)').get()
        if next_page:
            yield scrapy.Request(url=next_page, callback=self.parse)

    def parse_product(self, response):
        yield {
            'name': response.css('h2.product_title::text').get().strip(),
            'price': response.css('p.price span.woocommerce-Price-amount::text').get().strip()
        }

关键要点

  • Scrapy的异步架构不需要手动维护URL数组,通过yield scrapy.Request或CrawlSpider的Rules即可自动管理请求队列
  • 优先使用CSS/XPath选择器直接定位目标元素,避免遍历整个HTML,大幅提升解析效率
  • 如果分类页存在分页,必须添加分页处理逻辑,否则只能爬取第一页的产品

内容的提问来源于stack exchange,提问作者Bastion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 02:10:50