Scrapy爬虫爬取作者页面返回405错误,浏览器/Shell返回200
爬取作者页面返回405状态码的问题解决
问题背景
爬取http://quotes.toscrape.com/的作者数据时,运行爬虫出现大量405状态码,但浏览器或Scrapy Shell中请求相同URL却返回200。
原爬虫代码
import datetime import scrapy class AuthorsSpider(scrapy.Spider): name = 'authors' allowed_domains = ['quotes.toscrape.com'] start_urls = ['http://quotes.toscrape.com/'] custom_settings = { 'CONCURRENT_REQUESTS': 50, 'DOWNLOAD_DELAY': 0.1, 'FEED_URI': f'output/authors_{datetime.datetime.today().strftime("%Y-%m-%d %H-%M-%S")}.csv', 'FEED_FORMAT': 'csv', 'FEED_EXPORTERS': {'csv': 'scrapy.exporters.CsvItemExporter'}, 'FEED_EXPORT_ENCODING': 'utf-8', 'FEED_EXPORT_FIELDS': ('name','birth_date','birth_location','description',) } def parse(self, response): for _ in response.xpath("//div[@class='quote']"): author_page = response.xpath("//a[text()='(about)']/@href").get() yield response.follow(author_page, method="GET", callback=self.parse_author) next_page = response.xpath("//li[@class='next']/a/@href").get() if next_page: yield response.follow(next_page, self.parse) def parse_author(self, response): yield { 'name': response.xpath("//h3[@class='author-title']/text()").get(), 'birth_date': response.xpath("//span[@class='author-born-date']/text()").get(), 'birth_location': response.xpath("//span[@class='author-born-location']/text()").get(), 'description': response.xpath("//div[@class='author-description']/text()").get() }
运行日志片段
2023-01-02 10:53:33 [scrapy.core.engine] DEBUG: Crawled (200) <GET http://quotes.toscrape.com/page/10/> (referer: http://quotes.toscrape.com/page/9/) 2023-01-02 10:53:33 [scrapy.core.engine] DEBUG: Crawled (405) <NONE http://quotes.toscrape.com/author/Suzanne-Collins/> (referer: http://quotes.toscrape.com/page/7/) 2023-01-02 10:53:34 [scrapy.spidermiddlewares.httperror] INFO: Ignoring response <405 http://quotes.toscrape.com/author/Suzanne-Collins/>: HTTP status code is not handled or not allowed 2023-01-02 10:53:34 [scrapy.core.engine] DEBUG: Crawled (405) <NONE http://quotes.toscrape.com/author/W-C-Fields/> (referer: http://quotes.toscrape.com/page/8/) 2023-01-02 10:53:34 [scrapy.downloadermiddlewares.redirect] DEBUG: Redirecting (308) to <NONE http://quotes.toscrape.com/author/John-Lennon/> from <GET http://quotes.toscrape.com/author/John-Lennon> 2023-01-02 10:53:34 [scrapy.spidermiddlewares.httperror] INFO: Ignoring response <405 http://quotes.toscrape.com/author/W-C-Fields/>: HTTP status code is not handled or not allowed 2023-01-02 10:53:34 [scrapy.core.engine] DEBUG: Crawled (405) <NONE http://quotes.toscrape.com/author/Alfred-Tennyson/> (referer: http://quotes.toscrape.com/page/8/)
问题原因
- 重复请求同一作者链接:循环中使用全局XPath提取
(about)链接,每次都返回页面第一个作者的URL,导致单个页面内重复请求同一链接多次,触发服务器反爬限制,返回405。 - 高并发+低延迟放大风险:
CONCURRENT_REQUESTS=50并发过高,DOWNLOAD_DELAY=0.1延迟过低,对服务器造成较大压力,进一步触发反爬机制。 - 链接重定向导致异常:部分作者链接无尾部斜杠,服务器返回308重定向,重复请求叠加重定向导致Scrapy处理异常,状态码显示为405。
修复方案
1. 正确提取每个quote对应的作者链接
将循环内的XPath改为相对当前quote节点的查找,避免重复提取同一链接:
# 原代码 author_page = response.xpath("//a[text()='(about)']/@href").get() # 修改为 author_page = quote.xpath(".//a[text()='(about)']/@href").get()
注意开头的.,表示基于当前quote节点进行相对路径查找。
2. 调整并发与延迟参数
降低并发数,提高下载延迟,减少服务器压力:
'CONCURRENT_REQUESTS': 10, 'DOWNLOAD_DELAY': 0.5,
3. 移除冗余配置+优化数据提取
response.follow默认使用GET请求,可移除method="GET";- 对提取的文本添加
strip(),去除多余空格和换行; - 添加非空判断,避免无效请求。
修复后的完整代码
import datetime import scrapy class AuthorsSpider(scrapy.Spider): name = 'authors' allowed_domains = ['quotes.toscrape.com'] start_urls = ['http://quotes.toscrape.com/'] custom_settings = { 'CONCURRENT_REQUESTS': 10, 'DOWNLOAD_DELAY': 0.5, 'FEED_URI': f'output/authors_{datetime.datetime.today().strftime("%Y-%m-%d %H-%M-%S")}.csv', 'FEED_FORMAT': 'csv', 'FEED_EXPORTERS': {'csv': 'scrapy.exporters.CsvItemExporter'}, 'FEED_EXPORT_ENCODING': 'utf-8', 'FEED_EXPORT_FIELDS': ('name','birth_date','birth_location','description',) } def parse(self, response): for quote in response.xpath("//div[@class='quote']"): author_page = quote.xpath(".//a[text()='(about)']/@href").get() if author_page: yield response.follow(author_page, callback=self.parse_author) next_page = response.xpath("//li[@class='next']/a/@href").get() if next_page: yield response.follow(next_page, self.parse) def parse_author(self, response): yield { 'name': response.xpath("//h3[@class='author-title']/text()").get().strip(), 'birth_date': response.xpath("//span[@class='author-born-date']/text()").get(), 'birth_location': response.xpath("//span[@class='author-born-location']/text()").get(), 'description': response.xpath("//div[@class='author-description']/text()").get().strip() }
内容的提问来源于stack exchange,提问作者Fateme Fouladkar
相关产品推荐
相关产品推荐

