Python Requests.get请求纳斯达克北欧ETF历史价格页面时卡住的问题求助
Hey there, let's work through why your requests.get call is hanging even after adding a User-Agent. Sites like Nasdaq OMX Nordic usually have anti-scraping safeguards that require more than just a basic User-Agent to bypass. Here are a few fixes to try:
1. Add More Realistic Request Headers
Browsers send a lot more headers than just User-Agent when visiting a page. Adding these extra headers can make your request look more like a legitimate browser visit:
import requests url = "http://www.nasdaqomxnordic.com/etp/etf/etfhistorical?languageId=3&Instrument=SSE500" headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/39.0.2171.95 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8", "Accept-Language": "en-US,en;q=0.5", "Referer": "http://www.nasdaqomxnordic.com/etp/etf", "Connection": "keep-alive", "Upgrade-Insecure-Requests": "1" } try: # Add a timeout to avoid infinite hanging response = requests.get(url, headers=headers, timeout=15) # Check if the request was successful response.raise_for_status() # Verify by printing the first 500 characters of the response print("Response snippet:\n", response.text[:500]) except requests.exceptions.RequestException as e: print(f"Request failed with error: {str(e)}")
2. Check for Dynamic Content Loading
The page you're targeting might load the price data via an AJAX request after the initial page loads. Instead of scraping the main HTML page, you can directly call the underlying API:
- Open your browser's Developer Tools (F12)
- Go to the Network tab, then refresh the page
- Look for XHR/Fetch requests (usually named something like
etfhistoricaldataor similar) - Copy that API URL and use it in your
requestscall—this is often faster and more reliable than scraping the full page.
3. Use a Session to Maintain Cookies
Some sites require cookies to track session context. Using requests.Session() can help persist cookies across requests, making your session look more consistent:
import requests url = "http://www.nasdaqomxnordic.com/etp/etf/etfhistorical?languageId=3&Instrument=SSE500" headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/39.0.2171.95 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8" } session = requests.Session() session.headers.update(headers) try: # First make a dummy request to get necessary cookies session.get("http://www.nasdaqomxnordic.com/etp/etf", timeout=10) # Now request the historical data page response = session.get(url, timeout=15) response.raise_for_status() print("Response snippet:\n", response.text[:500]) except requests.exceptions.RequestException as e: print(f"Request failed with error: {str(e)}")
If none of these work, it's possible the site is using more advanced anti-scraping (like IP blocking or JavaScript checks). In that case, you might need to use a tool like Selenium to fully simulate a browser, but start with the above steps first—they usually solve the hanging request issue for this type of site.
内容的提问来源于stack exchange,提问作者JEB

