Python Requests+Scrapoxy无法获取代理添加的x-cache-proxyname头问题
Alright, let's break down exactly what's causing this flaky behavior and how to lock in a reliable fix:
- First, the
req.historylist in Python Requests only stores responses from redirects (like 3xx status codes that Requests automatically follows). If your target URL returns a 200 OK directly with no redirects,req.historywill be completely empty—so trying to accessreq.history[0]will throw that index-out-of-range error every single time. - Scrapoxy's
x-cache-proxynameheader isn't showing up consistently inhistorybecause it might only attach this header to redirect responses in some scenarios. When there's no redirect, the header gets added to the final response (the mainreqobject itself) instead of the history list. Your original code was probably only checkinghistory, hence the random failures.
Fixes to Implement
Fix 1: Check the Final Response First, Fall Back to History
This is the simplest, most robust approach—we'll prioritize checking the main response for the header, then iterate through the history if it's not found there:
import requests # Configure your Scrapoxy proxy details proxies = { "http": "http://your-scrapoxy-proxy-address:port", "https": "http://your-scrapoxy-proxy-address:port" } response = requests.get("https://your-target-url.com", proxies=proxies) # First check the final response's headers proxy_name = response.headers.get("x-cache-proxyname") # If not found, loop through each redirect response in history if not proxy_name: for redirect_resp in response.history: proxy_name = redirect_resp.headers.get("x-cache-proxyname") if proxy_name: break # Use the proxy name if we successfully retrieved it if proxy_name: print(f"Proxy used: {proxy_name}") else: print("Could not retrieve the proxy name header")
Fix 2: Disable Auto-Redirects & Handle Them Manually
If you need to capture the header reliably across every step of a redirect chain, turn off Requests' auto-follow behavior and handle each redirect yourself. This lets you inspect every single response that passes through Scrapoxy:
import requests proxies = { "http": "http://your-scrapoxy-proxy-address:port", "https": "http://your-scrapoxy-proxy-address:port" } target_url = "https://your-target-url.com" proxy_name = None while True: # Make the request without auto-following redirects response = requests.get(target_url, proxies=proxies, allow_redirects=False) # Check for the header in this specific response current_proxy = response.headers.get("x-cache-proxyname") if current_proxy: proxy_name = current_proxy break # If it's a valid redirect, update the target URL and continue if response.status_code in (301, 302, 303, 307, 308): target_url = response.headers.get("Location") if not target_url: break # No redirect URL found, exit the loop else: break # No redirect, exit the loop if proxy_name: print(f"Proxy used: {proxy_name}") else: print("Could not retrieve the proxy name header")
Bonus: Verify Scrapoxy's Configuration
Take a quick look at your Scrapoxy settings to ensure it's configured to inject the x-cache-proxyname header into all responses, not just redirects. Some proxy tools have conditional rules for header injection, so confirming this will eliminate any configuration-related gaps.
内容的提问来源于stack exchange,提问作者Filip Lav Maksimovic

