为何使用Python Requests库爬取https://mobile.lowvig.ag/sports返回403状态码(浏览器可正常访问)?求解决方法
Let's break down why you're hitting that 403 and how to fix it. Servers often return 403 when they detect non-browser traffic, so we need to make our Requests call mimic a real browser as closely as possible. Here's what to try:
1. Use a Full Set of Browser-like Headers
Your initial code doesn't send any custom headers, which is a huge red flag for anti-scraping systems. Even if you tried adding headers before, make sure you're including all critical ones that your browser sends.
Here's an updated code snippet with common headers matching a typical Chrome request:
import requests url = "https://mobile.lowvig.ag/sports" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8", "Accept-Language": "en-US,en;q=0.5", "Accept-Encoding": "gzip, deflate, br", "Connection": "keep-alive", "Upgrade-Insecure-Requests": "1", "Sec-Fetch-Dest": "document", "Sec-Fetch-Mode": "navigate", "Sec-Fetch-Site": "none", "Sec-Fetch-User": "?1" } response = requests.get(url, headers=headers) print(response.status_code) print(response.text[:500]) # Print first 500 chars to verify content
Why this works:
- User-Agent: Identifies your request as a real Chrome browser instead of the default Requests agent.
- Accept headers: Tell the server what content types/encodings your "browser" can handle, matching real browser behavior.
- Sec-Fetch headers: Modern browsers send these to indicate the request context (e.g., navigating to a document), helping the server validate legitimacy.
2. Use a Session to Maintain Cookies
Some sites require initial cookies (even for anonymous visits) before allowing access. Using requests.Session() automatically handles cookie persistence between requests, just like a browser does:
import requests url = "https://mobile.lowvig.ag/sports" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", # Include other headers from the snippet above } session = requests.Session() # First, make a dummy request to the homepage to grab initial cookies session.get("https://mobile.lowvig.ag", headers=headers) # Then fetch the sports page response = session.get(url, headers=headers) print(response.status_code)
3. Manually Add Browser Cookies (If Needed)
If the above still fails, inspect the cookies your browser sends when accessing the site:
- Open Chrome DevTools (F12) → Network tab.
- Refresh the page, click the first request to
mobile.lowvig.ag/sports. - Copy the
Cookieheader value from the "Request Headers" section.
Add it to your headers:
headers = { "User-Agent": "your-browser-user-agent", "Cookie": "paste-the-cookie-value-from-devtools", # Other headers... }
Note: Cookies expire quickly, so you'll need to refresh this value periodically.
4. When Requests Isn't Enough
If none of the above works, the site might use JavaScript to render content or have stricter anti-scraping measures (like Cloudflare). In that case, you'd need tools that execute JavaScript, such as Selenium or Playwright. But stick with the Requests-specific fixes first—they solve most basic 403 issues.
内容的提问来源于stack exchange,提问作者Ethan

