Scrapy爬虫请求目标URL返回403状态码的问题咨询
Hey there! Let's break down exactly why your Scrapy spider is hitting that 403 error, and walk through actionable fixes to get it working.
Common Causes of the 403 Error
- Unidentified User-Agent: By default, Scrapy sends a user-agent string like
Scrapy/2.8.0 (+https://scrapy.org)which makes it obvious this is a crawler. Most websites block these to prevent automated scraping. - Missing Request Headers: Browsers send a full set of headers (like
Accept,Accept-Language) with every request. Your current code doesn't include these, so the server can tell it's not a real user browsing. - Basic Anti-Scraping Checks: The site might have simple defenses in place, like checking for valid cookie context or requiring a small delay between requests (even for the first hit).
Step-by-Step Fixes
1. Add a Realistic User-Agent
The easiest first fix is to mimic a browser's user-agent. You can do this two ways:
Option A: Set Globally in settings.py
Open your project's settings.py file and update the USER_AGENT line:
USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
This applies the user-agent to all spiders in your project.
Option B: Set Per Spider (Custom Headers)
If you want to customize headers just for this spider, rewrite the start_requests method to include a full set of browser-like headers:
import scrapy class UsSpider(scrapy.Spider): name = 'us_spider' def start_requests(self): url = 'https://publicholidays.com/us/school-holidays/' # Mimic a Chrome browser's request headers headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Accept-Encoding': 'gzip, deflate, br', 'Connection': 'keep-alive' } yield scrapy.Request(url, headers=headers, callback=self.parse) def parse(self, response): print(response) print(response.request.headers) print("\n") yield { "hi": "hello" }
2. Add a Small Download Delay
Even if it's your first request, some sites flag rapid, consecutive requests. Add a delay in settings.py to make your crawler act more human:
DOWNLOAD_DELAY = 2 # Wait 2 seconds between requests
3. Handle Cookies (If Needed)
Some sites require valid cookies to serve content. You can enable Scrapy's built-in CookieMiddleware (it's enabled by default, but double-check in settings.py that it's not commented out):
DOWNLOADER_MIDDLEWARES = { 'scrapy.downloadermiddlewares.cookies.CookiesMiddleware': 700, }
If the site sets cookies on initial load, this middleware will automatically store and reuse them for subsequent requests.
4. Advanced: Check for JavaScript Rendering (If Basic Fixes Fail)
If you still get a 403 after trying the above, the site might be using JavaScript to load content or verify users. In that case, you can use tools like scrapy-playwright to render the page like a real browser.
Final Notes
Start with the user-agent and header fixes first—those resolve 90% of 403 issues with simple sites. If you're still blocked, check if your IP has been temporarily restricted (try switching networks or using a proxy).
内容的提问来源于stack exchange,提问作者Abu Horain

