使用Python urllib无法访问指定网站的技术求助
Hey Marco, I’ve dealt with whoscored.com’s anti-scraping measures a few times—they don’t just check for a basic User-Agent. Let’s walk through the most effective fixes for your issue:
1. Add More Comprehensive Request Headers
Your current code only sets the User-Agent, but modern sites verify a whole set of headers to confirm you’re a real browser. Update your headers to match what a typical Chrome request sends:
from urllib.request import Request, urlopen import urllib.error my_url = "https://www.whoscored.com/Statistics" headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.whoscored.com/', 'DNT': '1', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' } req = Request(my_url, headers=headers) try: page = urlopen(req).read() print("Success! Page content retrieved.") except urllib.error.HTTPError as e: print(f"HTTP Error: {e.code} - {e.reason}") except Exception as e: print(f"Error: {str(e)}")
The extra headers (like Referer and Accept) make your request look far more legitimate to whoscored’s servers.
2. Switch to the Requests Library (Easier for Session Management)
urllib works, but the requests library simplifies handling cookies, sessions, and headers—all critical for bypassing anti-scraping checks. First install it if you haven’t:
pip install requests
Then use a persistent session to mimic a real user’s browsing session:
import requests session = requests.Session() headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.whoscored.com/', 'DNT': '1' } # First visit the homepage to get necessary cookies session.get('https://www.whoscored.com/', headers=headers) # Now request the statistics page try: response = session.get('https://www.whoscored.com/Statistics', headers=headers) response.raise_for_status() # Raise error for HTTP status codes >=400 print("Success! Page content retrieved.") # Access content with response.text except requests.exceptions.HTTPError as e: print(f"HTTP Error: {e.response.status_code} - {e.response.reason}") except Exception as e: print(f"Error: {str(e)}")
Sessions automatically handle cookies, which many sites use to track valid users vs. scrapers.
3. Use a Proxy IP (If You’re IP-Blocked)
If whoscored has flagged your IP address, routing your request through a proxy can help. Here’s how to add a proxy to your requests code:
import requests proxy = { 'http': 'http://your-proxy-ip:port', 'https': 'http://your-proxy-ip:port' } session = requests.Session() session.proxies.update(proxy) # Rest of the code (headers, session.get) stays the same as above
Note: Free proxies are often unreliable—consider a paid service if you need consistent access.
4. Handle Dynamic Content (If Needed)
If the statistics page loads data dynamically with JavaScript, neither urllib nor requests will capture it directly. In that case, use Selenium to control a real browser:
from selenium import webdriver from selenium.webdriver.chrome.options import Options options = Options() options.add_argument('--headless=new') # Run in background without a window options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') driver = webdriver.Chrome(options=options) driver.get('https://www.whoscored.com/Statistics') # Wait for the page to load (adjust time as needed) driver.implicitly_wait(10) page_content = driver.page_source print(page_content) driver.quit()
You’ll need to install Selenium and the ChromeDriver matching your browser version for this to work.
Start with the first two solutions—they’re the quickest fixes for most cases with whoscored.com. If those don’t work, move to proxies or Selenium.
内容的提问来源于stack exchange,提问作者Marco Sousa-Poza

