使用Python 3.6.2爬取Google Finance遇urllib.error.HTTPError 403求助
Fixing HTTP Error 403: Forbidden When Scraping Google Finance with Python 3.6.2
Hey David, that 403 error is Google's anti-scraping mechanism flagging your request as coming from a script instead of a real browser—let's fix that quickly.
Why You're Getting This Error
By default, urllib.request sends a User-Agent header like Python-urllib/3.6, which is super easy for Google's systems to spot and block. They want to ensure traffic comes from legitimate users, not automated scrapers.
Quick Fix: Add Custom Request Headers
The simplest solution is to mimic a browser's request by adding a proper User-Agent header. Here's how to modify your code:
Modified Code
import urllib.request # Define the base URL (fixed your redundant url assignment here too!) url = "https://www.google.com/search?num=40&newwindow=1&tbm=fin&q=" stock_name = input("The stock you want to search for: ") url += stock_name # Mimic a Chrome browser request (update the User-Agent if needed) headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # Create a Request object with headers req = urllib.request.Request(url, headers=headers) try: data = urllib.request.urlopen(req).read() print("Success! Data retrieved.") # Process your data here except urllib.error.HTTPError as e: print(f"Error: {e}")
Key Changes Explained
- Added Headers: The
User-Agenttells Google your request is from a real Chrome browser. You can look up your own browser's User-Agent by searching "what's my user agent" if you want to use that instead. - Fixed Redundant URL Assignment: I removed the line where you overwrote
urlwith the fixedFBquery—now your user input forstock_nameactually gets used! - Error Handling: Wrapped the request in a
try/exceptblock to make error messages cleaner.
Additional Tips to Avoid Future Blocks
- Don't Request Too Frequently: Google will block your IP if you send too many requests in a short time. Add delays with
time.sleep()between requests. - Rotate User-Agents (If Needed): If you still get blocked after adding one User-Agent, try rotating between a few different browser User-Agents.
- Check Google's Terms of Service: Make sure your scraping activity complies with Google's terms—they restrict automated access to their services in most cases.
内容的提问来源于stack exchange,提问作者David Kong
相关产品推荐
相关产品推荐

