You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python urllib无法访问指定网站的技术求助

Fixing whoscored.com Blocking Your Python Request

Hey Marco, I’ve dealt with whoscored.com’s anti-scraping measures a few times—they don’t just check for a basic User-Agent. Let’s walk through the most effective fixes for your issue:

1. Add More Comprehensive Request Headers

Your current code only sets the User-Agent, but modern sites verify a whole set of headers to confirm you’re a real browser. Update your headers to match what a typical Chrome request sends:

from urllib.request import Request, urlopen
import urllib.error

my_url = "https://www.whoscored.com/Statistics"
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://www.whoscored.com/',
    'DNT': '1',
    'Connection': 'keep-alive',
    'Upgrade-Insecure-Requests': '1'
}
req = Request(my_url, headers=headers)
try:
    page = urlopen(req).read()
    print("Success! Page content retrieved.")
except urllib.error.HTTPError as e:
    print(f"HTTP Error: {e.code} - {e.reason}")
except Exception as e:
    print(f"Error: {str(e)}")

The extra headers (like Referer and Accept) make your request look far more legitimate to whoscored’s servers.

2. Switch to the Requests Library (Easier for Session Management)

urllib works, but the requests library simplifies handling cookies, sessions, and headers—all critical for bypassing anti-scraping checks. First install it if you haven’t:

pip install requests

Then use a persistent session to mimic a real user’s browsing session:

import requests

session = requests.Session()
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://www.whoscored.com/',
    'DNT': '1'
}

# First visit the homepage to get necessary cookies
session.get('https://www.whoscored.com/', headers=headers)

# Now request the statistics page
try:
    response = session.get('https://www.whoscored.com/Statistics', headers=headers)
    response.raise_for_status()  # Raise error for HTTP status codes >=400
    print("Success! Page content retrieved.")
    # Access content with response.text
except requests.exceptions.HTTPError as e:
    print(f"HTTP Error: {e.response.status_code} - {e.response.reason}")
except Exception as e:
    print(f"Error: {str(e)}")

Sessions automatically handle cookies, which many sites use to track valid users vs. scrapers.

3. Use a Proxy IP (If You’re IP-Blocked)

If whoscored has flagged your IP address, routing your request through a proxy can help. Here’s how to add a proxy to your requests code:

import requests

proxy = {
    'http': 'http://your-proxy-ip:port',
    'https': 'http://your-proxy-ip:port'
}

session = requests.Session()
session.proxies.update(proxy)
# Rest of the code (headers, session.get) stays the same as above

Note: Free proxies are often unreliable—consider a paid service if you need consistent access.

4. Handle Dynamic Content (If Needed)

If the statistics page loads data dynamically with JavaScript, neither urllib nor requests will capture it directly. In that case, use Selenium to control a real browser:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

options = Options()
options.add_argument('--headless=new')  # Run in background without a window
options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')

driver = webdriver.Chrome(options=options)
driver.get('https://www.whoscored.com/Statistics')

# Wait for the page to load (adjust time as needed)
driver.implicitly_wait(10)

page_content = driver.page_source
print(page_content)

driver.quit()

You’ll need to install Selenium and the ChromeDriver matching your browser version for this to work.

Start with the first two solutions—they’re the quickest fixes for most cases with whoscored.com. If those don’t work, move to proxies or Selenium.

内容的提问来源于stack exchange,提问作者Marco Sousa-Poza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:00:54