如何限制Scrapy CrawlSpider仅爬取目标网站前5页内容?
限制Scrapy CrawlSpider仅爬取前5页内容
问题背景
需使用Scrapy的CrawlSpider爬取文章的title、description字段,但要求仅爬取网站前5页内容,当前编写的CrawlSpider会爬取所有分页页面,需修改代码实现页数限制。
列表页HTML结构
<div class="list"> <div class="snippet-content"> <h2> <a href="https://example.com/article-1">Article 1</a> </h2> </div> <div class="snippet-content"> <h2> <a href="https://example.com/article-2">Article 2</a> </h2> </div> <div class="snippet-content"> <h2> <a href="https://example.com/article-3">Article 3</a> </h2> </div> <div class="snippet-content"> <h2> <a href="https://example.com/article-4">Article 4</a> </h2> </div> </div> <ul class="pagination"> <li class="next"> <a href="https://www.example.com?page=2&keywords=&from=&topic=&year=&type="> Next </a> </li> </ul>
详情页HTML结构
<meta property="og:title" content="Article Title"> <meta property="og:description" content="Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum has been the industry's standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing Lorem Ipsum passages, and more recently with desktop publishing software like Aldus PageMaker including versions of Lorem Ipsum.">
原CrawlSpider代码
from scrapy.linkextractors import LinkExtractor from scrapy.spiders import CrawlSpider, Rule import w3lib.html class ExampleSpider(CrawlSpider): name = "example" allowed_domains = ["example.com"] start_urls = ["https://www.example.com/"] custom_settings = { 'FEED_URI': 'articles.json', 'FEED_FORMAT': 'json' } total = 0 rules = ( # 提取当前页所有文章链接并跟进解析 Rule(LinkExtractor(restrict_xpaths='//div[contains(@class, "snippet-content")]/h2/a'), callback="parse_item", follow=True), # 提取下一页链接并自动跟进 Rule(LinkExtractor(restrict_xpaths='//ul[@class="pagination"]/li[@class="next"]/a')) ) def parse_item(self, response): self.total = self.total + 1 title = response.xpath('//meta[@property="og:title"]/@content').get() or "" description = w3lib.html.remove_tags(response.xpath('//meta[@property="og:description"]/@content').get()) or "" return { 'id': self.total, 'title': title, 'description': description }
解决方案
方案一:动态跟踪页码,限制分页跳转
通过解析URL中的page参数,判断当前页码是否小于设定的最大页数(5),仅满足条件时才继续跟进下一页。
修改后的代码:
from scrapy.linkextractors import LinkExtractor from scrapy.spiders import CrawlSpider, Rule import w3lib.html from urllib.parse import urlparse, parse_qs class ExampleSpider(CrawlSpider): name = "example" allowed_domains = ["example.com"] start_urls = ["https://www.example.com/"] custom_settings = { 'FEED_URI': 'articles.json', 'FEED_FORMAT': 'json' } total = 0 max_pages = 5 # 设定最大爬取页数 rules = ( # 提取文章详情页链接并解析 Rule(LinkExtractor(restrict_xpaths='//div[contains(@class, "snippet-content")]/h2/a'), callback="parse_item", follow=True), # 提取分页链接,用自定义回调处理是否跟进 Rule(LinkExtractor(restrict_xpaths='//ul[@class="pagination"]/li[@class="next"]/a'), callback="parse_pagination", follow=False), ) def parse_pagination(self, response): # 解析当前URL的page参数,默认值为1 parsed_url = urlparse(response.url) page_num = int(parse_qs(parsed_url.query).get('page', [1])[0]) # 仅当前页码小于最大页数时,才跟进下一页 if page_num < self.max_pages: yield response.follow(response.url) def parse_item(self, response): self.total = self.total + 1 title = response.xpath('//meta[@property="og:title"]/@content').get() or "" # og:description的content本身是纯文本,无需移除标签 description = response.xpath('//meta[@property="og:description"]/@content').get() or "" return { 'id': self.total, 'title': title, 'description': description }
修改要点:
- 添加
max_pages = 5常量,明确最大爬取页数 - 修改分页Rule,关闭自动跟进,改用自定义回调
parse_pagination处理分页逻辑 - 在
parse_pagination中解析当前页码,判断是否继续跳转 - 优化
description提取逻辑,移除不必要的w3lib.html.remove_tags调用
方案二:直接生成前5页URL作为起始链接
如果网站分页URL格式固定(如page=1到page=5),可直接生成所有目标页的URL作为start_urls,无需处理分页跳转。
修改后的代码:
from scrapy.linkextractors import LinkExtractor from scrapy.spiders import CrawlSpider, Rule import w3lib.html class ExampleSpider(CrawlSpider): name = "example" allowed_domains = ["example.com"] # 直接生成前5页的URL作为起始链接 start_urls = [f"https://www.example.com?page={page}" for page in range(1, 6)] custom_settings = { 'FEED_URI': 'articles.json', 'FEED_FORMAT': 'json' } total = 0 rules = ( # 仅提取文章详情页链接并解析,无需处理分页跳转 Rule(LinkExtractor(restrict_xpaths='//div[contains(@class, "snippet-content")]/h2/a'), callback="parse_item", follow=True), ) def parse_item(self, response): self.total = self.total + 1 title = response.xpath('//meta[@property="og:title"]/@content').get() or "" description = response.xpath('//meta[@property="og:description"]/@content').get() or "" return { 'id': self.total, 'title': title, 'description': description }
修改要点:
- 用列表推导式生成
page=1到page=5的URL作为start_urls - 删除分页相关的Rule,避免爬取更多分页
- 同样优化了
description的提取逻辑
内容的提问来源于stack exchange,提问作者Ven Nilson
相关产品推荐
相关产品推荐

