You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫爬取Google餐厅地址遇403禁止错误,请求帮助

Fixing Google Search Scraper: 403 Forbidden & Other Common Issues

Hey there! As someone who’s been through the early struggles of Python web scraping, let’s break down what’s going wrong with your crawler and fix it step by step.

1. Why You’re Getting the 403 Forbidden Error

Google actively blocks requests that look like they’re coming from bots. Your urllib request doesn’t include a User-Agent header—so Google immediately flags it as non-human traffic and rejects it. The fix is straightforward: add a valid User-Agent to mimic a real browser.

2. Other Bugs in Your Code

Let’s tackle the other issues in your script before we put it all together:

  • You forgot to import fromstring from lxml.html
  • The BeautifulSoup parser is set to 'url.parser' (invalid) — use 'lxml' or 'html.parser' instead
  • hotel_json is initialized as an empty dict, so accessing hotel_json["address"] will throw a KeyError
  • The variable html doesn’t exist (you meant the_page)
  • No check for whether the JSON-LD script actually contains an address field, which could cause crashes

Fixed & Annotated Code

Here’s the revised version with all fixes included:

import urllib.request, urllib.parse, urllib.error
from bs4 import BeautifulSoup
from lxml.html import fromstring  # Added missing import
import ssl
import json
import re
import sys
import warnings

# Ignore unnecessary warnings
if not sys.warnoptions:
    warnings.simplefilter("ignore")

# Add a valid User-Agent to mimic a browser request
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'
}

# Target URL (encoded properly)
url = "https://www.google.com/search?q=barbeque%20nation%20-%20noida"
request = urllib.request.Request(url, headers=headers)  # Include headers

try:
    response = urllib.request.urlopen(request)
    page = fromstring(response.read())  # Read response content first
    soup = BeautifulSoup(page, 'lxml')  # Use valid parser
    the_page = soup.prettify("utf-8")
    
    hotel_json = {"address": {}}  # Initialize address key to avoid KeyError
    
    # Loop through JSON-LD scripts to find restaurant details
    for line in soup.find_all('script', attrs={"type": "application/ld+json"}):
        try:
            details = json.loads(line.text.strip())
            # Check if the entry has address and streetAddress fields
            if "address" in details and "streetAddress" in details["address"]:
                hotel_json["address"]["LrzXr"] = details["address"]["streetAddress"]
                hotel_json["name"] = details.get("name", "Unknown_Restaurant")  # Fallback name if missing
                break
        except json.JSONDecodeError:
            continue  # Skip invalid JSON scripts
    
    # Save HTML and JSON files
    with open(f"{hotel_json['name']}.html", "wb") as file:
        file.write(the_page)  # Use correct variable name
    
    with open(f"{hotel_json['name']}.json", 'w') as outfile:
        json.dump(hotel_json, outfile, indent=4)
        
    print(f"Success! Saved data for {hotel_json['name']}")

except urllib.error.HTTPError as e:
    print(f"HTTP Error Occurred: {e}")
except Exception as e:
    print(f"Unexpected Error: {e}")

Key Changes Explained

  • User-Agent Header: Tells Google your request is coming from a real browser, bypassing the 403 block
  • Error Handling: Added try-except blocks to catch HTTP errors and invalid JSON
  • Proper Initialization: hotel_json now starts with an address key to prevent KeyErrors
  • Valid Parser: Switched to lxml for better HTML parsing (install it via pip install lxml if you haven’t)
  • Fallback Values: Added a backup name in case the JSON doesn’t include a restaurant name

A quick note: Google’s HTML structure can change over time, so this might need tweaks later. For more reliable results, consider using Google’s official Places API—it’s free for limited requests and avoids scraping headaches.

内容的提问来源于stack exchange,提问作者Anuj Pratap Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:07:13