使用Scrapy提取网页链接时仅处理首个匹配链接的问题排查
问题原因与解决办法
核心原因:Scrapy的
CrawlSpider规则里,一旦给Rule指定了callback参数,follow参数默认会设为False。这就导致爬虫爬完第一个匹配category/interviews的页面后,不会从该页面里继续提取符合规则的链接,直接停止了爬取。而如果不指定callback,follow默认是True,所以能正常遍历所有匹配规则的页面。解决办法:在你的
Rule里显式加上follow=True,这样即使指定了回调函数,爬虫也会继续从回调处理的页面中提取符合规则的链接,实现全量爬取。
修改后的代码如下:
from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor from scrapy.http import Request class MagazineCrawler(CrawlSpider): name = "MagazineCrawler" allowed_domains = ["eu-startups.com"] start_urls = ["https://www.eu-startups.com"] rules = ( Rule(LinkExtractor(allow=["category/interviews"]), callback="parse_category", follow=True), ) def parse_category(self, response): xpath_links = "//div[@class='td_block_inner tdb-block-inner td-fix-index']//a[@class='td-image-wrap ']/@href" subpage_links = response.xpath(xpath_links).extract() # Follow each subpage link and yield requests to crawl them for link in subpage_links: yield Request(link)
内容的提问来源于stack exchange,提问作者CoolMathematician
相关产品推荐
相关产品推荐

