stats.nba.com请求返回Response [403] 求排查问题原因
Hey Nick, let's break down why your code is hitting a 403 Forbidden error while other scripts for the same URL work—this almost always boils down to anti-scraping measures blocking your request. Here are the most likely fixes:
Common Causes & Fixes
1. Incomplete or Incorrect Request Headers
Stats.nba.com is strict about the headers it accepts, and missing key ones will trigger a 403. Your current code has a User-Agent list, but:
- The
desktop_agentslist looks truncated (ends with "Mozilla/5.0 (Windows ...")—make sure you have full, valid User-Agent strings. - You’re probably missing other critical headers like
Host,Referer,Accept, andAccept-Encoding.
Fix:
Grab the headers from your working script (or use your browser’s DevTools > Network tab to copy headers from a successful request) and add them to your code. Also, ensure you’re randomly picking a User-Agent for each request.
2. Missing Required Cookies or Request Parameters
Many NBA API endpoints require specific cookies (like __cf_bm, nba_session) or query parameters (like x-nba-stats-token) to validate requests. Your working script might already handle these automatically via a persistent session, while your current code doesn’t.
Fix:
- Use
requests.Session()instead of rawrequests.get()—this maintains cookies across requests, just like a browser does. - Check the Network tab of your working request to see if any unique parameters are being sent, and add them to your request URL or headers.
3. No Request Throttling
Even with random User-Agents, firing requests too quickly will trigger anti-scraping blocks. Your working script likely has delays between requests that yours is missing.
Fix:
Add random delays (1-3 seconds is a safe start) using time.sleep(uniform(1, 3)) to mimic human browsing speed.
4. Unverified SSL Certificates (Rare but Possible)
In some cases, requests might fail if SSL certificates aren’t verified, though this is less common for stats.nba.com.
Fix:
Add verify=False to your request (note: this is insecure for sensitive sites, but okay for public APIs like NBA stats) as a test, but prefer fixing the header/cookie issues first.
Modified Working Code Example
Here’s an updated version of your script incorporating all these fixes:
import requests import csv from random import choice, uniform import pandas as pd import time # Full, valid User-Agent list desktop_agents = [ 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/54.0.2840.99 Safari/537.36', 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/54.0.2840.99 Safari/537.36', 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' ] # Use a session to persist cookies session = requests.Session() def fetch_nba_stats(target_url): # Build complete headers headers = { 'User-Agent': choice(desktop_agents), 'Host': 'stats.nba.com', 'Referer': 'https://www.nba.com/', 'Accept': 'application/json, text/plain, */*', 'Accept-Encoding': 'gzip, deflate, br', 'Connection': 'keep-alive' } # First, hit the main NBA site to get necessary cookies session.get('https://www.nba.com/', headers=headers) # Add random delay to avoid rate limiting time.sleep(uniform(1, 3)) # Send the actual stats request response = session.get(target_url, headers=headers) if response.status_code == 200: return response.json() else: print(f"Request failed with status code: {response.status_code}") return None # Example usage stats_url = "YOUR_TARGET_STATS_NBA_URL" stats_data = fetch_nba_stats(stats_url) if stats_data: # Parse and save data to CSV if 'resultSets' in stats_data: df = pd.DataFrame( stats_data['resultSets'][0]['rowSet'], columns=stats_data['resultSets'][0]['headers'] ) df.to_csv('nba_stats.csv', index=False) print("Data saved successfully!")
Final Tip
Compare the exact request details (headers, cookies, parameters) between your failing script and the working one using tools like requests.utils.dict_from_cookiejar() or browser DevTools. The difference will almost always be the key to fixing the 403.
内容的提问来源于stack exchange,提问作者Nick

