You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy CrawlSpider嵌套爬取规则及分页处理技术问询

Scrapy CrawlSpider: Nested Rules & Pagination Handling

Hey there! Let's break this down for you since you're new to Scrapy's CrawlSpider—this is a super common scenario, and it's totally manageable once you understand how CrawlSpider's rules work.

First off: Yes, CrawlSpider absolutely supports the "nested" crawling flow you're aiming for (category pages → product pages + pagination). It doesn't use strict "hierarchy" labels, but we can use rule order, restrict_xpaths, and follow flags to replicate that logic perfectly.

Let's Fix Your Rules Step-by-Step

Your core issue is ordering and targeting the right links at the right time. Here's how to structure your rules correctly:

1. Understand Rule Priority

CrawlSpider processes rules in the order you define them. So we want to put the most specific rules first (like product links) before broader ones (like category pagination or new category pages). This prevents links from being caught by the wrong rule.

2. Final Rule Setup

Here's the adjusted rule set with comments explaining each part:

from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor

class YourSpider(CrawlSpider):
    name = "your_spider"
    allowed_domains = ["your-target-domain.com"]  # Replace with your site's domain
    start_urls = ["https://your-target-domain.com/"]

    rules = (
        # Rule 1: Catch product links on category pages
        # Target the <a> tags inside product items (adjust XPath if your site's structure differs)
        Rule(
            LinkExtractor(restrict_xpaths='//div[@class="sch-category-products-item"]/a'),
            callback="parse_product",
            follow=False  # No need to follow links from product pages (unless you need to scrape related products)
        ),

        # Rule 2: Follow pagination links on category pages
        # Replace the XPath with your actual pagination link selector (e.g., next/prev page buttons)
        Rule(
            LinkExtractor(restrict_xpaths='//div[contains(@class, "pagination")]/a'),
            follow=True  # Follow these links to load more product pages in the same category
        ),

        # Rule 3: Follow category links from the main menu
        Rule(
            LinkExtractor(restrict_xpaths='//div[@class="sch-main-menu-sub-links-left"]/a'),
            callback="parse_item",
            follow=True  # Follow into category pages to trigger Rules 1 & 2
        ),
    )

    def parse_item(self, response):
        # This runs when we first load a category page
        # Extract category-level data here (e.g., category name, description)
        category_data = {
            "category_name": response.xpath("//h1/text()").get().strip(),
            "category_url": response.url
        }
        yield category_data

    def parse_product(self, response):
        # This runs for every product page
        # Extract product details here
        product_data = {
            "product_name": response.xpath("//h2[@class='product-title']/text()").get().strip(),
            "product_price": response.xpath("//span[@class='product-price']/text()").get(),
            "product_url": response.url
        }
        yield product_data

Key Explanations

  • Rule Order: We first grab product links (most specific), then handle pagination (to load more products in the same category), finally follow new category links. This ensures products are scraped before moving to new categories or pages.
  • Pagination Handling: The pagination rule doesn't need a callback—we just set follow=True because when we load a paginated category page, Rule 1 will automatically kick in to scrape the new set of products.
  • XPath Tweaks: I adjusted your product link XPath to target the <a> tag directly (/a at the end) since LinkExtractor needs to find actual <a href="..."> elements to crawl. Double-check your pagination XPath to match your site's structure (e.g., if pagination uses a class like sch-pagination, update that selector).

How the Crawl Flow Works

  1. Start at your homepage → Rule 3 matches main menu category links → crawls into each category page, runs parse_item to save category data.
  2. On the category page → Rule 1 matches all product links → crawls each product page, runs parse_product to save product data.
  3. Also on the category page → Rule 2 matches pagination links → crawls the next page of the same category, repeating steps 2-3 until all pages are scraped.
  4. Once all pages/products in a category are done, the crawler moves back to the main menu to process the next category.

That's exactly the nested, paginated crawl you need!

内容的提问来源于stack exchange,提问作者F. Shahid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:41:00