You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何使用Python Requests库爬取https://mobile.lowvig.ag/sports返回403状态码(浏览器可正常访问)?求解决方法

Fixing 403 Forbidden When Scraping https://mobile.lowvig.ag/sports with Python Requests

Let's break down why you're hitting that 403 and how to fix it. Servers often return 403 when they detect non-browser traffic, so we need to make our Requests call mimic a real browser as closely as possible. Here's what to try:

1. Use a Full Set of Browser-like Headers

Your initial code doesn't send any custom headers, which is a huge red flag for anti-scraping systems. Even if you tried adding headers before, make sure you're including all critical ones that your browser sends.

Here's an updated code snippet with common headers matching a typical Chrome request:

import requests

url = "https://mobile.lowvig.ag/sports"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.5",
    "Accept-Encoding": "gzip, deflate, br",
    "Connection": "keep-alive",
    "Upgrade-Insecure-Requests": "1",
    "Sec-Fetch-Dest": "document",
    "Sec-Fetch-Mode": "navigate",
    "Sec-Fetch-Site": "none",
    "Sec-Fetch-User": "?1"
}

response = requests.get(url, headers=headers)
print(response.status_code)
print(response.text[:500])  # Print first 500 chars to verify content

Why this works:

  • User-Agent: Identifies your request as a real Chrome browser instead of the default Requests agent.
  • Accept headers: Tell the server what content types/encodings your "browser" can handle, matching real browser behavior.
  • Sec-Fetch headers: Modern browsers send these to indicate the request context (e.g., navigating to a document), helping the server validate legitimacy.

2. Use a Session to Maintain Cookies

Some sites require initial cookies (even for anonymous visits) before allowing access. Using requests.Session() automatically handles cookie persistence between requests, just like a browser does:

import requests

url = "https://mobile.lowvig.ag/sports"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    # Include other headers from the snippet above
}

session = requests.Session()
# First, make a dummy request to the homepage to grab initial cookies
session.get("https://mobile.lowvig.ag", headers=headers)
# Then fetch the sports page
response = session.get(url, headers=headers)

print(response.status_code)

3. Manually Add Browser Cookies (If Needed)

If the above still fails, inspect the cookies your browser sends when accessing the site:

  1. Open Chrome DevTools (F12) → Network tab.
  2. Refresh the page, click the first request to mobile.lowvig.ag/sports.
  3. Copy the Cookie header value from the "Request Headers" section.

Add it to your headers:

headers = {
    "User-Agent": "your-browser-user-agent",
    "Cookie": "paste-the-cookie-value-from-devtools",
    # Other headers...
}

Note: Cookies expire quickly, so you'll need to refresh this value periodically.

4. When Requests Isn't Enough

If none of the above works, the site might use JavaScript to render content or have stricter anti-scraping measures (like Cloudflare). In that case, you'd need tools that execute JavaScript, such as Selenium or Playwright. But stick with the Requests-specific fixes first—they solve most basic 403 issues.


内容的提问来源于stack exchange,提问作者Ethan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 23:54:21