Python爬虫爬取Google餐厅地址遇403禁止错误,请求帮助
Hey there! As someone who’s been through the early struggles of Python web scraping, let’s break down what’s going wrong with your crawler and fix it step by step.
1. Why You’re Getting the 403 Forbidden Error
Google actively blocks requests that look like they’re coming from bots. Your urllib request doesn’t include a User-Agent header—so Google immediately flags it as non-human traffic and rejects it. The fix is straightforward: add a valid User-Agent to mimic a real browser.
2. Other Bugs in Your Code
Let’s tackle the other issues in your script before we put it all together:
- You forgot to import
fromstringfromlxml.html - The BeautifulSoup parser is set to
'url.parser'(invalid) — use'lxml'or'html.parser'instead hotel_jsonis initialized as an empty dict, so accessinghotel_json["address"]will throw a KeyError- The variable
htmldoesn’t exist (you meantthe_page) - No check for whether the JSON-LD script actually contains an
addressfield, which could cause crashes
Fixed & Annotated Code
Here’s the revised version with all fixes included:
import urllib.request, urllib.parse, urllib.error from bs4 import BeautifulSoup from lxml.html import fromstring # Added missing import import ssl import json import re import sys import warnings # Ignore unnecessary warnings if not sys.warnoptions: warnings.simplefilter("ignore") # Add a valid User-Agent to mimic a browser request headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36' } # Target URL (encoded properly) url = "https://www.google.com/search?q=barbeque%20nation%20-%20noida" request = urllib.request.Request(url, headers=headers) # Include headers try: response = urllib.request.urlopen(request) page = fromstring(response.read()) # Read response content first soup = BeautifulSoup(page, 'lxml') # Use valid parser the_page = soup.prettify("utf-8") hotel_json = {"address": {}} # Initialize address key to avoid KeyError # Loop through JSON-LD scripts to find restaurant details for line in soup.find_all('script', attrs={"type": "application/ld+json"}): try: details = json.loads(line.text.strip()) # Check if the entry has address and streetAddress fields if "address" in details and "streetAddress" in details["address"]: hotel_json["address"]["LrzXr"] = details["address"]["streetAddress"] hotel_json["name"] = details.get("name", "Unknown_Restaurant") # Fallback name if missing break except json.JSONDecodeError: continue # Skip invalid JSON scripts # Save HTML and JSON files with open(f"{hotel_json['name']}.html", "wb") as file: file.write(the_page) # Use correct variable name with open(f"{hotel_json['name']}.json", 'w') as outfile: json.dump(hotel_json, outfile, indent=4) print(f"Success! Saved data for {hotel_json['name']}") except urllib.error.HTTPError as e: print(f"HTTP Error Occurred: {e}") except Exception as e: print(f"Unexpected Error: {e}")
Key Changes Explained
- User-Agent Header: Tells Google your request is coming from a real browser, bypassing the 403 block
- Error Handling: Added try-except blocks to catch HTTP errors and invalid JSON
- Proper Initialization:
hotel_jsonnow starts with anaddresskey to prevent KeyErrors - Valid Parser: Switched to
lxmlfor better HTML parsing (install it viapip install lxmlif you haven’t) - Fallback Values: Added a backup name in case the JSON doesn’t include a restaurant name
A quick note: Google’s HTML structure can change over time, so this might need tweaks later. For more reliable results, consider using Google’s official Places API—it’s free for limited requests and avoids scraping headaches.
内容的提问来源于stack exchange,提问作者Anuj Pratap Singh

