如何在Scrapy CrawlSpider中限制爬取页面数量?
如何在Scrapy CrawlSpider中限制爬取页面数量?
嘿,我来帮你搞定这个问题!要在Scrapy的CrawlSpider里只爬取前5页(哪怕网站有50页),最直接的方法是给爬虫加一个页面计数器,通过控制分页链接的跟进逻辑来实现。下面是具体的修改方案:
方法:添加分页计数器控制
我们可以在爬虫类里维护一个计数器,每次处理下一页链接时检查计数,达到5页就停止跟进后续分页。
修改后的完整代码如下:
from scrapy.linkextractors import LinkExtractor from scrapy.spiders import CrawlSpider, Rule class BooksSpider(CrawlSpider): name = "bookscraper" allowed_domains = ["books.toscrape.com"] start_urls = ["https://books.toscrape.com/"] def __init__(self, *args, **kwargs): super(BooksSpider, self).__init__(*args, **kwargs) self.page_count = 1 # 起始页算作第1页 def process_next_page_links(self, links): # 检查当前页数,达到5页就不再返回下一页链接 if self.page_count >= 5: return [] self.page_count += 1 return links rules = ( Rule(LinkExtractor(restrict_xpaths='//h3/a'), callback='parse_item', follow=True), # 给分页规则加上链接处理函数,控制分页数量 Rule(LinkExtractor(restrict_xpaths='//li[@class="next"]/a'), follow=True, process_links='process_next_page_links'), ) def parse_item(self, response): product_info = response.xpath('//table[contains(@class, "table-striped")]') name = response.xpath('//h1/text()').get() upc = product_info.xpath('(./tr/td)[1]/text()').get() price = product_info.xpath('(./tr/td)[3]/text()').get() availability = product_info.xpath('(./tr/td)[6]/text()').get() yield {'Name': name, 'UPC': upc, 'Availability': availability, 'Price': price}
原理说明
- 初始化时
self.page_count = 1,因为start_urls对应的是第1页; process_next_page_links函数会在每次提取到下一页链接时被调用:如果当前页数已经到5,就返回空列表,爬虫就不会再跟进下一页;否则计数器加1,正常返回链接继续爬取;- 这样就能精准控制只爬取前5页的内容啦!
备注:内容来源于stack exchange,提问作者Bilal Anees
相关产品推荐
相关产品推荐

