You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取SteamDB配置请求头后仍报403错误如何修复

Scrapy配置请求头后爬取SteamDB仍返回403错误的解决方案

问题概述

  • 爬取目标:SteamDB数据统计页面
  • 异常表现:已手动配置请求头的前提下,Scrapy发起的GET请求持续返回403状态码,响应被爬虫中间件拦截,无法进入parse方法执行解析逻辑。

复现代码

def start_request(self):
        headers =  {"user-agent": "Mozilla/5.0 (Linux; Android 6.0; Nexus 5 Build/MRA58N) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/101.0.4951.67 Mobile Safari/537.36",
"accept": "application/json",
"accept-encoding": "gzip, deflate, br",
"accept-language": "en-US,en;q=0.9,en-GB;q=0.8,ar;q=0.7",
"cache-control":" no-cache",
"pragma": "no-cache",
"referer": "https://steamdb.info/graph/", 
"sec-fetch-dest": "empty",
"sec-fetch-mode": "cors",
"sec-fetch-site": "same-origin",
"x-requested-with": "XMLHttpRequest"
            }

        yield scrapy.Request(url = 'https://steamdb.info/graph', method='GET', headers = headers, callback=self.parse)
        

    def parse(self, response):    
        # 解析逻辑

报错日志

2022-07-08 20:20:41 [scrapy.core.engine] DEBUG: Crawled (403) <GET https://steamdb.info/graph> (referer: https://steamdb.info/graph/)
2022-07-08 20:20:41 [scrapy.spidermiddlewares.httperror] INFO: Ignoring response <403 https://steamdb.info/graph>: HTTP status code is not handled or not allowed

问题根因

  • 请求头配置不符合真实浏览器访问逻辑:当前配置的accept: application/json、sec-fetch-dest: empty、x-requested-with: XMLHttpRequest属于页面内异步拉取数据的XHR请求特征,直接访问静态页面路径时携带这些头会直接触发反爬拦截。另外首次访问站点时sec-fetch-site值为none,配置为same-origin同时携带自引用referer,和真实用户首次进入页面的请求特征完全不符。
  • Scrapy原生请求的TLS握手特征、请求头排序规则和真实浏览器存在差异,即便请求头文本完全一致,也会被站点使用的Cloudflare反爬系统识别为自动化工具。
  • Scrapy默认仅放行200-299区间的HTTP响应到回调方法,未手动配置允许403状态码的前提下,拦截响应是默认行为。

修复方案

  1. 修正请求头配置,匹配真实浏览器访问静态页面的特征,删除异步请求专属头字段:
headers = {
    "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36",
    "accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
    "accept-encoding": "gzip, deflate, br",
    "accept-language": "en-US,en;q=0.9",
    "cache-control": "no-cache",
    "pragma": "no-cache",
    "sec-fetch-dest": "document",
    "sec-fetch-mode": "navigate",
    "sec-fetch-site": "none",
    "sec-fetch-user": "?1",
    "upgrade-insecure-requests": "1"
}

注意:首次访问不要携带referer,否则会被判定为异常请求

  1. 绕过TLS指纹校验:
    安装scrapy-playwright组件,调用真实Chromium内核发起请求,从底层模拟浏览器行为,安装命令:

    pip install scrapy-playwright
    playwright install chromium
    

    随后在settings.py中添加如下配置:

    DOWNLOAD_HANDLERS = {
        "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
        "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    }
    TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"
    ROBOTSTXT_OBEY = False
    DOWNLOAD_DELAY = 1.5
    

    发起请求时指定使用playwright加载:

    yield scrapy.Request(
        url='https://steamdb.info/graph',
        meta={"playwright": True},
        headers=headers,
        callback=self.parse
    )
    
  2. 若需要调试403响应内容,在爬虫类中添加配置,允许403状态码进入回调:

class SteamdbSpider(scrapy.Spider):
    name = "steamdb"
    allowed_domains = ["steamdb.info"]
    handle_httpstatus_list = [403]

内容的提问来源于stack exchange,提问作者Baraa Zaid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 23:54:18