You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python BeautifulSoup爬取Tripadvisor咖啡馆信息及数据清洗求助

Fixing Your TripAdvisor Web Scraping Code

Hey there! Let's get your Python BeautifulSoup script working properly so you can extract and format cafe info exactly how you want it. First, let's go over the key issues in your current code, then we'll rewrite it step by step.

Issues in Your Current Code

  • You forgot to import urllib.request, so your get_HTML function will throw an error right away.
  • Each function re-parses the entire HTML document (super inefficient!) and ignores the passed HTML parameter, relying on a global variable instead.
  • Your selectors are too rigid (like only targeting 4.5-star scores) and don't link cafe names, ratings, and review counts together—you risk ending up with mismatched data.
  • You're printing raw elements instead of cleaning and formatting the text into your desired output style.

Improved Working Code

from bs4 import BeautifulSoup
import urllib.request

def get_html(url):
    # Add a user-agent header to avoid being blocked by TripAdvisor
    headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'}
    request = urllib.request.Request(url, headers=headers)
    response = urllib.request.urlopen(request)
    return response.read()

def extract_cafe_info(html):
    soup = BeautifulSoup(html, "lxml")
    # Target the main container for each cafe entry (adjust class if TripAdvisor updates its structure)
    cafe_cards = soup.find_all('div', class_='biGQs _P pZUbB osNWb')
    
    for card in cafe_cards:
        # Extract and clean cafe name
        cafe_name = card.find('a', class_='fHibz').get_text(strip=True)
        
        # Extract rating from the image's alt attribute
        rating_img = card.find('img', class_='zWXXY')
        rating_text = rating_img['alt'] if rating_img else 'No rating available'
        
        # Extract and clean review count
        review_count_elem = card.find('span', class_='IiChw')
        review_count = review_count_elem.get_text(strip=True).split()[0] if review_count_elem else '0'
        
        # Format output exactly as requested
        print(f"{cafe_name}, {rating_text}, {review_count} reviews.")

# Run the script
target_url = 'https://www.tripadvisor.com.au/Restaurants-g255068-c8-Brisbane_Brisbane_Region_Queensland.html'
tripadvisor_html = get_html(target_url)
extract_cafe_info(tripadvisor_html)

Key Improvements Explained

  • User-Agent Header: TripAdvisor blocks requests without a proper user-agent, so this prevents 403 Forbidden errors.
  • Single HTML Parsing: We parse the HTML once instead of three times, making the script faster and more efficient.
  • Linked Data Extraction: We grab each cafe's entire card first, so all info (name, rating, reviews) comes from the same entry—no mismatched data.
  • Error Handling: Added checks for missing elements (like cafes with no ratings) to prevent the script from crashing.
  • Cleaned Text: Used get_text(strip=True) to remove extra whitespace, and formatted output to match your requested style perfectly.

Quick Tips for Future Scraping

  • TripAdvisor’s HTML structure can change over time. If the script stops working, use your browser’s DevTools to inspect the page and update the class names.
  • For larger scraping tasks, consider using the requests library instead of urllib (it’s more user-friendly) and add small delays between requests to avoid being blocked.

内容的提问来源于stack exchange,提问作者noob123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:41:10