使用Python的mechanize库访问网站时出现超时错误求助
Hey there, I’ve run into similar headaches with NSE’s website before—they’ve got pretty strict anti-scraping measures that flag basic mechanize requests as bots right away. Let’s walk through some fixes that should get your code working:
1. Spoof a Real Browser User-Agent
Mechanize uses a generic default user-agent that’s dead easy for NSE’s systems to spot. Add a genuine browser UA to your request to blend in:
import mechanize url = 'https://www1.nseindia.com/products/content/equities/equities/eq_security.htm' br = mechanize.Browser() br.set_handle_refresh(False) br.set_handle_equiv(False) # Use a real user-agent (update this to match your current browser's UA if needed) br.addheaders = [('User-Agent', 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')] try: # Set a longer timeout (30 seconds) to account for slow server responses response = br.open(url, timeout=30) print("Success! Response status:", response.code) except mechanize.HTTPError as e: print(f"HTTP Error: {e.code} - {e.reason}") except mechanize.URLError as e: print(f"URL Error: {e.reason}")
2. Fix Redirect Handling
NSE often redirects traffic between subdomains (like from www1.nseindia.com to nseindia.com). Tweak mechanize’s redirect settings to follow these properly:
br.set_handle_redirect(True) # Let mechanize automatically follow redirects br.set_handle_refresh(mechanize._http.HTTPRefreshProcessor(), max_time=1)
3. Try a More Modern Alternative (Requests)
Mechanize is a bit outdated these days—many developers switch to requests (paired with BeautifulSoup if you need to parse content) because it handles anti-scraping measures more smoothly. Here’s a quick replacement:
import requests url = 'https://www1.nseindia.com/products/content/equities/equities/eq_security.htm' headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9' } try: response = requests.get(url, headers=headers, timeout=30, allow_redirects=True) response.raise_for_status() # Trigger error for bad status codes (like 403/404) print("Success! Response status:", response.status_code) except requests.exceptions.RequestException as e: print(f"Request failed: {e}")
4. Quick Extra Tips
- If timeouts persist, check your network—NSE’s servers get slow during market hours.
- Avoid spamming requests; add a small delay (
time.sleep(2)) between calls to mimic human behavior. - Mechanize handles cookies by default, but you can verify if any are being set with
print(br.cookiejar)—some parts of NSE’s site require valid cookies to load.
内容的提问来源于stack exchange,提问作者Vinu

