You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫爬取作者页面返回405错误,浏览器/Shell返回200

爬取作者页面返回405状态码的问题解决

问题背景

爬取http://quotes.toscrape.com/的作者数据时,运行爬虫出现大量405状态码,但浏览器或Scrapy Shell中请求相同URL却返回200。

原爬虫代码

import datetime
import scrapy

class AuthorsSpider(scrapy.Spider):
    name = 'authors'
    allowed_domains = ['quotes.toscrape.com']
    start_urls = ['http://quotes.toscrape.com/']
    custom_settings = {
        'CONCURRENT_REQUESTS': 50,
        'DOWNLOAD_DELAY': 0.1,
        'FEED_URI': f'output/authors_{datetime.datetime.today().strftime("%Y-%m-%d %H-%M-%S")}.csv',
        'FEED_FORMAT': 'csv',
        'FEED_EXPORTERS': {'csv': 'scrapy.exporters.CsvItemExporter'},
        'FEED_EXPORT_ENCODING': 'utf-8',
        'FEED_EXPORT_FIELDS': ('name','birth_date','birth_location','description',) 
    }

    def parse(self, response):
        for _ in response.xpath("//div[@class='quote']"):
            author_page = response.xpath("//a[text()='(about)']/@href").get()
            yield response.follow(author_page,
                                method="GET",
                                callback=self.parse_author)

        next_page = response.xpath("//li[@class='next']/a/@href").get()
        if next_page:
            yield response.follow(next_page, self.parse)


    def parse_author(self, response):
        yield {
            'name': response.xpath("//h3[@class='author-title']/text()").get(),
            'birth_date': response.xpath("//span[@class='author-born-date']/text()").get(),
            'birth_location': response.xpath("//span[@class='author-born-location']/text()").get(),
            'description': response.xpath("//div[@class='author-description']/text()").get()
        }

运行日志片段

2023-01-02 10:53:33 [scrapy.core.engine] DEBUG: Crawled (200) <GET http://quotes.toscrape.com/page/10/> (referer: http://quotes.toscrape.com/page/9/)
2023-01-02 10:53:33 [scrapy.core.engine] DEBUG: Crawled (405) <NONE http://quotes.toscrape.com/author/Suzanne-Collins/> (referer: http://quotes.toscrape.com/page/7/)
2023-01-02 10:53:34 [scrapy.spidermiddlewares.httperror] INFO: Ignoring response <405 http://quotes.toscrape.com/author/Suzanne-Collins/>: HTTP status code is not handled or not allowed
2023-01-02 10:53:34 [scrapy.core.engine] DEBUG: Crawled (405) <NONE http://quotes.toscrape.com/author/W-C-Fields/> (referer: http://quotes.toscrape.com/page/8/)
2023-01-02 10:53:34 [scrapy.downloadermiddlewares.redirect] DEBUG: Redirecting (308) to <NONE http://quotes.toscrape.com/author/John-Lennon/> from <GET http://quotes.toscrape.com/author/John-Lennon>
2023-01-02 10:53:34 [scrapy.spidermiddlewares.httperror] INFO: Ignoring response <405 http://quotes.toscrape.com/author/W-C-Fields/>: HTTP status code is not handled or not allowed
2023-01-02 10:53:34 [scrapy.core.engine] DEBUG: Crawled (405) <NONE http://quotes.toscrape.com/author/Alfred-Tennyson/> (referer: http://quotes.toscrape.com/page/8/)

问题原因

  1. 重复请求同一作者链接:循环中使用全局XPath提取(about)链接,每次都返回页面第一个作者的URL,导致单个页面内重复请求同一链接多次,触发服务器反爬限制,返回405。
  2. 高并发+低延迟放大风险:CONCURRENT_REQUESTS=50并发过高,DOWNLOAD_DELAY=0.1延迟过低,对服务器造成较大压力,进一步触发反爬机制。
  3. 链接重定向导致异常:部分作者链接无尾部斜杠,服务器返回308重定向,重复请求叠加重定向导致Scrapy处理异常,状态码显示为405。

修复方案

1. 正确提取每个quote对应的作者链接

将循环内的XPath改为相对当前quote节点的查找,避免重复提取同一链接:

# 原代码
author_page = response.xpath("//a[text()='(about)']/@href").get()
# 修改为
author_page = quote.xpath(".//a[text()='(about)']/@href").get()

注意开头的.,表示基于当前quote节点进行相对路径查找。

2. 调整并发与延迟参数

降低并发数,提高下载延迟,减少服务器压力:

'CONCURRENT_REQUESTS': 10,
'DOWNLOAD_DELAY': 0.5,

3. 移除冗余配置+优化数据提取

  • response.follow默认使用GET请求,可移除method="GET";
  • 对提取的文本添加strip(),去除多余空格和换行;
  • 添加非空判断,避免无效请求。

修复后的完整代码

import datetime
import scrapy

class AuthorsSpider(scrapy.Spider):
    name = 'authors'
    allowed_domains = ['quotes.toscrape.com']
    start_urls = ['http://quotes.toscrape.com/']
    custom_settings = {
        'CONCURRENT_REQUESTS': 10,
        'DOWNLOAD_DELAY': 0.5,
        'FEED_URI': f'output/authors_{datetime.datetime.today().strftime("%Y-%m-%d %H-%M-%S")}.csv',
        'FEED_FORMAT': 'csv',
        'FEED_EXPORTERS': {'csv': 'scrapy.exporters.CsvItemExporter'},
        'FEED_EXPORT_ENCODING': 'utf-8',
        'FEED_EXPORT_FIELDS': ('name','birth_date','birth_location','description',) 
    }

    def parse(self, response):
        for quote in response.xpath("//div[@class='quote']"):
            author_page = quote.xpath(".//a[text()='(about)']/@href").get()
            if author_page:
                yield response.follow(author_page, callback=self.parse_author)

        next_page = response.xpath("//li[@class='next']/a/@href").get()
        if next_page:
            yield response.follow(next_page, self.parse)


    def parse_author(self, response):
        yield {
            'name': response.xpath("//h3[@class='author-title']/text()").get().strip(),
            'birth_date': response.xpath("//span[@class='author-born-date']/text()").get(),
            'birth_location': response.xpath("//span[@class='author-born-location']/text()").get(),
            'description': response.xpath("//div[@class='author-description']/text()").get().strip()
        }

内容的提问来源于stack exchange,提问作者Fateme Fouladkar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 08:35:24