You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫请求目标URL返回403状态码的问题咨询

Why You're Getting a 403 Forbidden Error & How to Fix It

Hey there! Let's break down exactly why your Scrapy spider is hitting that 403 error, and walk through actionable fixes to get it working.

Common Causes of the 403 Error

  • Unidentified User-Agent: By default, Scrapy sends a user-agent string like Scrapy/2.8.0 (+https://scrapy.org) which makes it obvious this is a crawler. Most websites block these to prevent automated scraping.
  • Missing Request Headers: Browsers send a full set of headers (like Accept, Accept-Language) with every request. Your current code doesn't include these, so the server can tell it's not a real user browsing.
  • Basic Anti-Scraping Checks: The site might have simple defenses in place, like checking for valid cookie context or requiring a small delay between requests (even for the first hit).

Step-by-Step Fixes

1. Add a Realistic User-Agent

The easiest first fix is to mimic a browser's user-agent. You can do this two ways:

Option A: Set Globally in settings.py

Open your project's settings.py file and update the USER_AGENT line:

USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'

This applies the user-agent to all spiders in your project.

Option B: Set Per Spider (Custom Headers)

If you want to customize headers just for this spider, rewrite the start_requests method to include a full set of browser-like headers:

import scrapy

class UsSpider(scrapy.Spider):
    name = 'us_spider'

    def start_requests(self):
        url = 'https://publicholidays.com/us/school-holidays/'
        # Mimic a Chrome browser's request headers
        headers = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
            'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
            'Accept-Language': 'en-US,en;q=0.5',
            'Accept-Encoding': 'gzip, deflate, br',
            'Connection': 'keep-alive'
        }
        yield scrapy.Request(url, headers=headers, callback=self.parse)

    def parse(self, response):
        print(response)
        print(response.request.headers)
        print("\n")
        yield { "hi": "hello" }

2. Add a Small Download Delay

Even if it's your first request, some sites flag rapid, consecutive requests. Add a delay in settings.py to make your crawler act more human:

DOWNLOAD_DELAY = 2  # Wait 2 seconds between requests

3. Handle Cookies (If Needed)

Some sites require valid cookies to serve content. You can enable Scrapy's built-in CookieMiddleware (it's enabled by default, but double-check in settings.py that it's not commented out):

DOWNLOADER_MIDDLEWARES = {
    'scrapy.downloadermiddlewares.cookies.CookiesMiddleware': 700,
}

If the site sets cookies on initial load, this middleware will automatically store and reuse them for subsequent requests.

4. Advanced: Check for JavaScript Rendering (If Basic Fixes Fail)

If you still get a 403 after trying the above, the site might be using JavaScript to load content or verify users. In that case, you can use tools like scrapy-playwright to render the page like a real browser.

Final Notes

Start with the user-agent and header fixes first—those resolve 90% of 403 issues with simple sites. If you're still blocked, check if your IP has been temporarily restricted (try switching networks or using a proxy).

内容的提问来源于stack exchange,提问作者Abu Horain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 13:13:16