You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫无法爬取下一页,请求技术协助

Scrapy爬虫无法爬取下页的问题解决

问题现象

爬完第一页数据后,日志出现两次DEBUG: Crawled (200) <GET https://reedsy.com/robots.txt> (referer: None),随后提示:

[scrapy.dupefilters] DEBUG: Filtered duplicate request: <GET https://reedsy.com/> - no more duplicates will be shown (see DUPEFILTER_DEBUG to show all duplicates)

原爬虫代码

import scrapy

class PublisherSpider(scrapy.Spider):
    name = 'mycrawler'
    start_urls = ['https://blog.reedsy.com/publishers/african-american/?accepts_submissions=true&formats=&publisher_size=']
   
    def parse(self, response):
        for publishers in response.css('div.panel-body'):
            publisher = publishers.css('h3.text-heavy::text').get()
            url = publishers.css('a.text-blue::attr(href)').get()
            if publisher and url:
                yield {"Publisher": publisher.strip(), "url": url}
                
        next_page = response.css('a::attr(href)').get()
        if next_page is not None:
            yield response.follow(next_page, callback = self.parse)

曾尝试的写法

next_page = response.css('a').attrib['href']
yield response.follow(next_page, callback = self.parse, dont_filter = True)
next_page = response.css('a::attr(href)').extract()
next_page = response.css('a::attr(href)').extract_first() 

问题根源与解决方法

问题核心

你用的a::attr(href)选择器太宽泛,会匹配页面上所有a标签,第一个匹配到的大概率是网站首页的链接(比如导航栏logo),导致爬虫跳转到首页而非下一页,重复请求首页触发了去重过滤。

正确做法

  1. 定位下一页按钮的精准选择器:
    打开目标页面,按F12打开浏览器开发者工具,找到分页栏里的「Next」按钮,查看它的专属class或属性。比如该网站下一页按钮通常带rel="next"属性,或者有pagination-next这类class。

  2. 修改代码里的下一页获取逻辑:
    替换成精准的选择器,比如:

    # 示例1:匹配带rel="next"的下一页按钮
     next_page = response.css('a[rel="next"]::attr(href)').get()
     if next_page:
         yield response.follow(next_page, callback=self.parse)
    

    或者如果按钮有专属class:

    # 示例2:匹配带有pagination-next类的按钮
     next_page = response.css('a.pagination-next::attr(href)').get()
     if next_page:
         yield response.follow(next_page, callback=self.parse)
    

注意事项

  • 不要随便加dont_filter=True,这会导致大量重复请求,浪费资源还可能触发反爬机制。
  • response.follow会自动处理相对URL,无需手动拼接绝对地址。
  • 一定要用浏览器开发者工具验证选择器,确保只匹配到下一页链接,避免无关链接干扰。

内容的提问来源于stack exchange,提问作者TLit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 17:27:25