Scrapy爬虫无法爬取下一页,请求技术协助
Scrapy爬虫无法爬取下页的问题解决
问题现象
爬完第一页数据后,日志出现两次DEBUG: Crawled (200) <GET https://reedsy.com/robots.txt> (referer: None),随后提示:
[scrapy.dupefilters] DEBUG: Filtered duplicate request: <GET https://reedsy.com/> - no more duplicates will be shown (see DUPEFILTER_DEBUG to show all duplicates)
原爬虫代码
import scrapy class PublisherSpider(scrapy.Spider): name = 'mycrawler' start_urls = ['https://blog.reedsy.com/publishers/african-american/?accepts_submissions=true&formats=&publisher_size='] def parse(self, response): for publishers in response.css('div.panel-body'): publisher = publishers.css('h3.text-heavy::text').get() url = publishers.css('a.text-blue::attr(href)').get() if publisher and url: yield {"Publisher": publisher.strip(), "url": url} next_page = response.css('a::attr(href)').get() if next_page is not None: yield response.follow(next_page, callback = self.parse)
曾尝试的写法
next_page = response.css('a').attrib['href'] yield response.follow(next_page, callback = self.parse, dont_filter = True) next_page = response.css('a::attr(href)').extract() next_page = response.css('a::attr(href)').extract_first()
问题根源与解决方法
问题核心
你用的a::attr(href)选择器太宽泛,会匹配页面上所有a标签,第一个匹配到的大概率是网站首页的链接(比如导航栏logo),导致爬虫跳转到首页而非下一页,重复请求首页触发了去重过滤。
正确做法
定位下一页按钮的精准选择器:
打开目标页面,按F12打开浏览器开发者工具,找到分页栏里的「Next」按钮,查看它的专属class或属性。比如该网站下一页按钮通常带rel="next"属性,或者有pagination-next这类class。修改代码里的下一页获取逻辑:
替换成精准的选择器,比如:# 示例1:匹配带rel="next"的下一页按钮 next_page = response.css('a[rel="next"]::attr(href)').get() if next_page: yield response.follow(next_page, callback=self.parse)或者如果按钮有专属class:
# 示例2:匹配带有pagination-next类的按钮 next_page = response.css('a.pagination-next::attr(href)').get() if next_page: yield response.follow(next_page, callback=self.parse)
注意事项
- 不要随便加
dont_filter=True,这会导致大量重复请求,浪费资源还可能触发反爬机制。 response.follow会自动处理相对URL,无需手动拼接绝对地址。- 一定要用浏览器开发者工具验证选择器,确保只匹配到下一页链接,避免无关链接干扰。
内容的提问来源于stack exchange,提问作者TLit
相关产品推荐
相关产品推荐

