Python入门4个月:requests.get请求部分URL无响应问题求助
requests.get() Calls Hey there! Let's break down why some URLs hang when using requests.get() and how to fix them without just catching exceptions—requests absolutely can handle these scenarios, you just need to tweak your approach a bit.
Fix 1: Spoof Request Headers to Bypass Anti-Scraping Blocks
Most websites flag the default requests user agent (like python-requests/2.31.0) as non-human traffic and block it. Mimicking a browser's request headers is the easiest fix, starting with the User-Agent field. Adding extra fields like Accept-Language can help make your request look more legitimate.
Example code:
import requests headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9' } # Test with your problematic URL response = requests.get("https://your-problem-url.com", headers=headers, timeout=10) print(response.status_code)
Fix 2: Use a Session to Maintain Context
Some websites require persistent session cookies (e.g., for session validation or tracking) to serve content. Using requests.Session() automatically retains cookies between requests, which can resolve hangs caused by missing context.
Example code:
import requests # Create a session to keep cookies session = requests.Session() # First visit the site's homepage to get necessary session cookies session.get("https://target-site.com") # Now request the problematic URL within the same session response = session.get("https://your-problem-url.com", timeout=10) print(response.text[:500]) # Print first 500 chars of response
Fix 3: Resolve SSL Certificate Issues
If the target website has an untrusted or self-signed SSL certificate, requests may stall during certificate verification (even with a timeout set). You can either disable verification (note: this has security risks) or specify a trusted certificate file.
Example code (disable verification temporarily):
response = requests.get("https://your-problem-url.com", verify=False, timeout=10)
Fix 4: Try a Proxy (If IP is Restricted)
If your IP address is blocked by the target website, using a proxy can bypass this restriction. Find a reliable HTTP/HTTPS proxy and pass it to the proxies parameter.
Example code:
proxies = { 'http': 'http://your-proxy-server:port', 'https': 'https://your-proxy-server:port' } response = requests.get("https://your-problem-url.com", proxies=proxies, timeout=10)
Final Note
Most unresponsive requests.get() calls stem from the target website rejecting your request for being "non-standard," not a limitation of the requests library. Start with adjusting request headers and using a session—these fixes resolve 80% of common issues. If you still hit walls, check the website's robots.txt to ensure you're allowed to scrape it, or inspect the browser's network tab to see if the site requires specific POST parameters or other request configurations.
内容的提问来源于stack exchange,提问作者kerem guner

