使用Scrapy请求返回403错误,但requests.get可正常访问
我尝试用Scrapy抓取多个网站内容,但所有目标网站都返回403(Forbidden)状态码。不过用requests库的get方法能正常访问这些网站,代码示例如下:
import requests url = "https://www.name_of_website.com/" headers = { "User-Agent": "Mozilla/5.0 (compatible; MSIE 9.0; Windows NT 6.1; WOW64; Trident/5.0)", } response = requests.get(url, headers=headers) print(response.status_code)
另外,这些网站用Chrome浏览器也能正常访问。我已经在Scrapy的DEFAULT_REQUEST_HEADERS里配置了和Chrome相同的请求头,但还是失败。
我搞不懂为什么Scrapy请求失败,而requests.get()能正常工作,这个问题在多个网站上都出现了。我还试过用scrapy-fake-useragent及相关中间件,也没效果。之前看过类似问题但没得到帮助,现在求新思路。
补充测试信息
出于研究目的,我测试了以下网址,响应状态码如下:
https://www.fastcompany.com/ - 403 https://www.ft.com/ - 200 https://www.theinformation.com/ - 200 https://www.pcmag.com/ - 403 https://www.thestreet.com/ - 403
这些网址都无法通过Scrapy正常访问。
我的Scrapy爬虫代码
class TheinformationSpider(scrapy.Spider): name = "theinformation" allowed_domains = ["www.theinformation.com"] start_urls = ["https://www.theinformation.com/"] def parse(self, response): print(response)
目前我只关注响应状态码。
更新后的Scrapy设置
DEFAULT_REQUEST_HEADERS = { "User-Agent": "Mozilla/5.0 (Linux; x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/121.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7", "Accept-Encoding": "gzip, deflate, br, zstd", "Accept-Language": "en-US,en;q=0.9", "Referer": "http://www.google.com", }
爬取时的响应信息
2024-03-08 15:15:54 [scrapy.core.engine] DEBUG: Crawled (403) <GET https://www.theinformation.com/> (referer: http://www.google.com) 2024-03-08 15:15:54 [scrapy.spidermiddlewares.httperror] INFO: Ignoring response <403 https://www.theinformation.com/>: HTTP status code is not handled or not allowed 2024-03-08 15:15:54 [scrapy.core.engine] INFO: Closing spider (finished) Total articles scrapped by "theinformation" = 0, null data = 0
解决思路
1. 逐字段对比请求头差异
虽然配置了DEFAULT_REQUEST_HEADERS,但Scrapy可能自动添加额外请求头(比如Connection、Host),或者对Accept-Encoding的处理逻辑和requests不一致。用抓包工具(Charles、Fiddler)分别抓取Scrapy、requests、浏览器的完整请求,对比所有请求头字段,重点排查:
Connection字段是否和浏览器一致(通常为keep-alive)- 浏览器访问时生成的Cookie,Scrapy初始请求是否缺失
Cache-Control这类容易被忽略的标识字段
2. 复用requests的会话Cookie
很多网站会在首次访问时设置反爬相关Cookie,Scrapy直接请求会缺少这些验证信息。可以先用requests建立会话获取Cookie,再传给Scrapy:
import requests from scrapy.http import Request class TheinformationSpider(scrapy.Spider): name = "theinformation" allowed_domains = ["www.theinformation.com"] start_urls = ["https://www.theinformation.com/"] def start_requests(self): url = self.start_urls[0] headers = { "User-Agent": "Mozilla/5.0 (Linux; x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/121.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7", "Accept-Language": "en-US,en;q=0.9", "Referer": "http://www.google.com", } # 用requests会话获取Cookie session = requests.Session() session.get(url, headers=headers) cookies = session.cookies.get_dict() # 携带Cookie发起Scrapy请求 yield Request(url, cookies=cookies, headers=headers) def parse(self, response): print(response.status_code)
3. 调整Scrapy下载器设置
- 关闭
HttpCompressionMiddleware或移除Accept-Encoding中的zstd部分:部分服务器不支持该编码格式,可能触发反爬 - 设置
DOWNLOAD_DELAY = 2,模拟人类访问的间隔 - 禁用可能添加额外标识的中间件,确保请求头完全自定义
4. 启用JS渲染中间件
如果网站依赖JS生成验证参数(如Cloudflare验证),Scrapy的纯HTTP请求会被拦截。可以使用scrapy-playwright模拟浏览器渲染:
# settings.py配置 DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } PLAYWRIGHT_LAUNCH_OPTIONS = {"headless": True} # 爬虫代码 class TheinformationSpider(scrapy.Spider): name = "theinformation" allowed_domains = ["www.theinformation.com"] start_urls = ["https://www.theinformation.com/"] def start_requests(self): for url in self.start_urls: yield scrapy.Request(url, meta={"playwright": True}) def parse(self, response): print(response.status_code)
5. 验证IP是否被标记
即使requests能访问,Scrapy的请求特征可能被网站识别为爬虫。可以配置代理IP:
# settings.py DOWNLOADER_MIDDLEWARES = { 'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware': 110, } PROXIES = ['http://your-proxy-ip:port']
内容的提问来源于stack exchange,提问作者Mah3sh

