Scrapy CrawlSpider嵌套爬取规则及分页处理技术问询
Hey there! Let's break this down for you since you're new to Scrapy's CrawlSpider—this is a super common scenario, and it's totally manageable once you understand how CrawlSpider's rules work.
First off: Yes, CrawlSpider absolutely supports the "nested" crawling flow you're aiming for (category pages → product pages + pagination). It doesn't use strict "hierarchy" labels, but we can use rule order, restrict_xpaths, and follow flags to replicate that logic perfectly.
Let's Fix Your Rules Step-by-Step
Your core issue is ordering and targeting the right links at the right time. Here's how to structure your rules correctly:
1. Understand Rule Priority
CrawlSpider processes rules in the order you define them. So we want to put the most specific rules first (like product links) before broader ones (like category pagination or new category pages). This prevents links from being caught by the wrong rule.
2. Final Rule Setup
Here's the adjusted rule set with comments explaining each part:
from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor class YourSpider(CrawlSpider): name = "your_spider" allowed_domains = ["your-target-domain.com"] # Replace with your site's domain start_urls = ["https://your-target-domain.com/"] rules = ( # Rule 1: Catch product links on category pages # Target the <a> tags inside product items (adjust XPath if your site's structure differs) Rule( LinkExtractor(restrict_xpaths='//div[@class="sch-category-products-item"]/a'), callback="parse_product", follow=False # No need to follow links from product pages (unless you need to scrape related products) ), # Rule 2: Follow pagination links on category pages # Replace the XPath with your actual pagination link selector (e.g., next/prev page buttons) Rule( LinkExtractor(restrict_xpaths='//div[contains(@class, "pagination")]/a'), follow=True # Follow these links to load more product pages in the same category ), # Rule 3: Follow category links from the main menu Rule( LinkExtractor(restrict_xpaths='//div[@class="sch-main-menu-sub-links-left"]/a'), callback="parse_item", follow=True # Follow into category pages to trigger Rules 1 & 2 ), ) def parse_item(self, response): # This runs when we first load a category page # Extract category-level data here (e.g., category name, description) category_data = { "category_name": response.xpath("//h1/text()").get().strip(), "category_url": response.url } yield category_data def parse_product(self, response): # This runs for every product page # Extract product details here product_data = { "product_name": response.xpath("//h2[@class='product-title']/text()").get().strip(), "product_price": response.xpath("//span[@class='product-price']/text()").get(), "product_url": response.url } yield product_data
Key Explanations
- Rule Order: We first grab product links (most specific), then handle pagination (to load more products in the same category), finally follow new category links. This ensures products are scraped before moving to new categories or pages.
- Pagination Handling: The pagination rule doesn't need a callback—we just set
follow=Truebecause when we load a paginated category page, Rule 1 will automatically kick in to scrape the new set of products. - XPath Tweaks: I adjusted your product link XPath to target the
<a>tag directly (/aat the end) since LinkExtractor needs to find actual<a href="...">elements to crawl. Double-check your pagination XPath to match your site's structure (e.g., if pagination uses a class likesch-pagination, update that selector).
How the Crawl Flow Works
- Start at your homepage → Rule 3 matches main menu category links → crawls into each category page, runs
parse_itemto save category data. - On the category page → Rule 1 matches all product links → crawls each product page, runs
parse_productto save product data. - Also on the category page → Rule 2 matches pagination links → crawls the next page of the same category, repeating steps 2-3 until all pages are scraped.
- Once all pages/products in a category are done, the crawler moves back to the main menu to process the next category.
That's exactly the nested, paginated crawl you need!
内容的提问来源于stack exchange,提问作者F. Shahid

