You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何限制Scrapy CrawlSpider仅爬取目标网站前5页内容?

限制Scrapy CrawlSpider仅爬取前5页内容

问题背景

需使用Scrapy的CrawlSpider爬取文章的title、description字段,但要求仅爬取网站前5页内容,当前编写的CrawlSpider会爬取所有分页页面,需修改代码实现页数限制。

列表页HTML结构

<div class="list">
  <div class="snippet-content">
    <h2>
      <a href="https://example.com/article-1">Article 1</a>
    </h2>
  </div>
  <div class="snippet-content">
    <h2>
      <a href="https://example.com/article-2">Article 2</a>
    </h2>
  </div>
  <div class="snippet-content">
    <h2>
      <a href="https://example.com/article-3">Article 3</a>
    </h2>
  </div>
  <div class="snippet-content">
    <h2>
      <a href="https://example.com/article-4">Article 4</a>
    </h2>
  </div>
</div>
<ul class="pagination">
  <li class="next">
    <a href="https://www.example.com?page=2&keywords=&from=&topic=&year=&type="> Next </a>
  </li>
</ul>

详情页HTML结构

<meta property="og:title" content="Article Title">
<meta property="og:description" content="Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum has been the industry's standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing Lorem Ipsum passages, and more recently with desktop publishing software like Aldus PageMaker including versions of Lorem Ipsum.">

原CrawlSpider代码

from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule
import w3lib.html


class ExampleSpider(CrawlSpider):
    name = "example"
    allowed_domains = ["example.com"]
    start_urls = ["https://www.example.com/"]
    custom_settings = {
        'FEED_URI': 'articles.json',
        'FEED_FORMAT': 'json'
    }
    total = 0

   
    rules = (
        # 提取当前页所有文章链接并跟进解析
        Rule(LinkExtractor(restrict_xpaths='//div[contains(@class, "snippet-content")]/h2/a'), callback="parse_item",
             follow=True),
        # 提取下一页链接并自动跟进
        Rule(LinkExtractor(restrict_xpaths='//ul[@class="pagination"]/li[@class="next"]/a'))
    )

    def parse_item(self, response):
        self.total = self.total + 1
        title = response.xpath('//meta[@property="og:title"]/@content').get() or ""
        description = w3lib.html.remove_tags(response.xpath('//meta[@property="og:description"]/@content').get()) or ""
       
        return {
            'id': self.total,
            'title': title,
            'description': description
        }

解决方案

方案一:动态跟踪页码,限制分页跳转

通过解析URL中的page参数,判断当前页码是否小于设定的最大页数(5),仅满足条件时才继续跟进下一页。

修改后的代码:

from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule
import w3lib.html
from urllib.parse import urlparse, parse_qs


class ExampleSpider(CrawlSpider):
    name = "example"
    allowed_domains = ["example.com"]
    start_urls = ["https://www.example.com/"]
    custom_settings = {
        'FEED_URI': 'articles.json',
        'FEED_FORMAT': 'json'
    }
    total = 0
    max_pages = 5  # 设定最大爬取页数

    rules = (
        # 提取文章详情页链接并解析
        Rule(LinkExtractor(restrict_xpaths='//div[contains(@class, "snippet-content")]/h2/a'), 
             callback="parse_item", follow=True),
        # 提取分页链接,用自定义回调处理是否跟进
        Rule(LinkExtractor(restrict_xpaths='//ul[@class="pagination"]/li[@class="next"]/a'), 
             callback="parse_pagination", follow=False),
    )

    def parse_pagination(self, response):
        # 解析当前URL的page参数,默认值为1
        parsed_url = urlparse(response.url)
        page_num = int(parse_qs(parsed_url.query).get('page', [1])[0])
        
        # 仅当前页码小于最大页数时,才跟进下一页
        if page_num < self.max_pages:
            yield response.follow(response.url)

    def parse_item(self, response):
        self.total = self.total + 1
        title = response.xpath('//meta[@property="og:title"]/@content').get() or ""
        # og:description的content本身是纯文本,无需移除标签
        description = response.xpath('//meta[@property="og:description"]/@content').get() or ""
        
        return {
            'id': self.total,
            'title': title,
            'description': description
        }

修改要点:

  • 添加max_pages = 5常量,明确最大爬取页数
  • 修改分页Rule,关闭自动跟进,改用自定义回调parse_pagination处理分页逻辑
  • 在parse_pagination中解析当前页码,判断是否继续跳转
  • 优化description提取逻辑,移除不必要的w3lib.html.remove_tags调用

方案二:直接生成前5页URL作为起始链接

如果网站分页URL格式固定(如page=1到page=5),可直接生成所有目标页的URL作为start_urls,无需处理分页跳转。

修改后的代码:

from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule
import w3lib.html


class ExampleSpider(CrawlSpider):
    name = "example"
    allowed_domains = ["example.com"]
    # 直接生成前5页的URL作为起始链接
    start_urls = [f"https://www.example.com?page={page}" for page in range(1, 6)]
    custom_settings = {
        'FEED_URI': 'articles.json',
        'FEED_FORMAT': 'json'
    }
    total = 0

    rules = (
        # 仅提取文章详情页链接并解析,无需处理分页跳转
        Rule(LinkExtractor(restrict_xpaths='//div[contains(@class, "snippet-content")]/h2/a'), 
             callback="parse_item", follow=True),
    )

    def parse_item(self, response):
        self.total = self.total + 1
        title = response.xpath('//meta[@property="og:title"]/@content').get() or ""
        description = response.xpath('//meta[@property="og:description"]/@content').get() or ""
        
        return {
            'id': self.total,
            'title': title,
            'description': description
        }

修改要点:

  • 用列表推导式生成page=1到page=5的URL作为start_urls
  • 删除分页相关的Rule,避免爬取更多分页
  • 同样优化了description的提取逻辑

内容的提问来源于stack exchange,提问作者Ven Nilson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 04:15:03