You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy请求返回403错误,但requests.get可正常访问

Scrapy请求返回403但requests和浏览器可正常访问的问题

我尝试用Scrapy抓取多个网站内容,但所有目标网站都返回403(Forbidden)状态码。不过用requests库的get方法能正常访问这些网站,代码示例如下:

import requests
url = "https://www.name_of_website.com/"
headers = {
    "User-Agent": "Mozilla/5.0 (compatible; MSIE 9.0; Windows NT 6.1; WOW64; Trident/5.0)",
}
response = requests.get(url, headers=headers)
print(response.status_code)

另外,这些网站用Chrome浏览器也能正常访问。我已经在Scrapy的DEFAULT_REQUEST_HEADERS里配置了和Chrome相同的请求头,但还是失败。

我搞不懂为什么Scrapy请求失败,而requests.get()能正常工作,这个问题在多个网站上都出现了。我还试过用scrapy-fake-useragent及相关中间件,也没效果。之前看过类似问题但没得到帮助,现在求新思路。

补充测试信息

出于研究目的,我测试了以下网址,响应状态码如下:

https://www.fastcompany.com/        -  403
https://www.ft.com/                 -  200
https://www.theinformation.com/     -  200
https://www.pcmag.com/              -  403
https://www.thestreet.com/          -  403

这些网址都无法通过Scrapy正常访问。

我的Scrapy爬虫代码

class TheinformationSpider(scrapy.Spider):
    name = "theinformation"
    allowed_domains = ["www.theinformation.com"]
    start_urls = ["https://www.theinformation.com/"]

    def parse(self, response):
       print(response)

目前我只关注响应状态码。

更新后的Scrapy设置

DEFAULT_REQUEST_HEADERS = {
    "User-Agent": "Mozilla/5.0 (Linux; x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/121.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7",
    "Accept-Encoding": "gzip, deflate, br, zstd",
    "Accept-Language": "en-US,en;q=0.9",
    "Referer": "http://www.google.com",
}

爬取时的响应信息

2024-03-08 15:15:54 [scrapy.core.engine] DEBUG: Crawled (403) <GET https://www.theinformation.com/> (referer: http://www.google.com)
2024-03-08 15:15:54 [scrapy.spidermiddlewares.httperror] INFO: Ignoring response <403 https://www.theinformation.com/>: HTTP status code is not handled or not allowed
2024-03-08 15:15:54 [scrapy.core.engine] INFO: Closing spider (finished)
Total articles scrapped by "theinformation" = 0, null data = 0

解决思路

1. 逐字段对比请求头差异

虽然配置了DEFAULT_REQUEST_HEADERS,但Scrapy可能自动添加额外请求头(比如Connection、Host),或者对Accept-Encoding的处理逻辑和requests不一致。用抓包工具(Charles、Fiddler)分别抓取Scrapy、requests、浏览器的完整请求,对比所有请求头字段,重点排查:

  • Connection字段是否和浏览器一致(通常为keep-alive)
  • 浏览器访问时生成的Cookie,Scrapy初始请求是否缺失
  • Cache-Control这类容易被忽略的标识字段

2. 复用requests的会话Cookie

很多网站会在首次访问时设置反爬相关Cookie,Scrapy直接请求会缺少这些验证信息。可以先用requests建立会话获取Cookie,再传给Scrapy:

import requests
from scrapy.http import Request

class TheinformationSpider(scrapy.Spider):
    name = "theinformation"
    allowed_domains = ["www.theinformation.com"]
    start_urls = ["https://www.theinformation.com/"]

    def start_requests(self):
        url = self.start_urls[0]
        headers = {
            "User-Agent": "Mozilla/5.0 (Linux; x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/121.0.0.0 Safari/537.36",
            "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7",
            "Accept-Language": "en-US,en;q=0.9",
            "Referer": "http://www.google.com",
        }
        # 用requests会话获取Cookie
        session = requests.Session()
        session.get(url, headers=headers)
        cookies = session.cookies.get_dict()
        # 携带Cookie发起Scrapy请求
        yield Request(url, cookies=cookies, headers=headers)

    def parse(self, response):
       print(response.status_code)

3. 调整Scrapy下载器设置

  • 关闭HttpCompressionMiddleware或移除Accept-Encoding中的zstd部分:部分服务器不支持该编码格式,可能触发反爬
  • 设置DOWNLOAD_DELAY = 2,模拟人类访问的间隔
  • 禁用可能添加额外标识的中间件,确保请求头完全自定义

4. 启用JS渲染中间件

如果网站依赖JS生成验证参数(如Cloudflare验证),Scrapy的纯HTTP请求会被拦截。可以使用scrapy-playwright模拟浏览器渲染:

# settings.py配置
DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}
PLAYWRIGHT_LAUNCH_OPTIONS = {"headless": True}

# 爬虫代码
class TheinformationSpider(scrapy.Spider):
    name = "theinformation"
    allowed_domains = ["www.theinformation.com"]
    start_urls = ["https://www.theinformation.com/"]

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(url, meta={"playwright": True})

    def parse(self, response):
       print(response.status_code)

5. 验证IP是否被标记

即使requests能访问,Scrapy的请求特征可能被网站识别为爬虫。可以配置代理IP:

# settings.py
DOWNLOADER_MIDDLEWARES = {
    'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware': 110,
}
PROXIES = ['http://your-proxy-ip:port']

内容的提问来源于stack exchange,提问作者Mah3sh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 11:37:04