Python BeautifulSoup爬取Tripadvisor咖啡馆信息及数据清洗求助
Fixing Your TripAdvisor Web Scraping Code
Hey there! Let's get your Python BeautifulSoup script working properly so you can extract and format cafe info exactly how you want it. First, let's go over the key issues in your current code, then we'll rewrite it step by step.
Issues in Your Current Code
- You forgot to import
urllib.request, so yourget_HTMLfunction will throw an error right away. - Each function re-parses the entire HTML document (super inefficient!) and ignores the passed
HTMLparameter, relying on a global variable instead. - Your selectors are too rigid (like only targeting 4.5-star scores) and don't link cafe names, ratings, and review counts together—you risk ending up with mismatched data.
- You're printing raw elements instead of cleaning and formatting the text into your desired output style.
Improved Working Code
from bs4 import BeautifulSoup import urllib.request def get_html(url): # Add a user-agent header to avoid being blocked by TripAdvisor headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} request = urllib.request.Request(url, headers=headers) response = urllib.request.urlopen(request) return response.read() def extract_cafe_info(html): soup = BeautifulSoup(html, "lxml") # Target the main container for each cafe entry (adjust class if TripAdvisor updates its structure) cafe_cards = soup.find_all('div', class_='biGQs _P pZUbB osNWb') for card in cafe_cards: # Extract and clean cafe name cafe_name = card.find('a', class_='fHibz').get_text(strip=True) # Extract rating from the image's alt attribute rating_img = card.find('img', class_='zWXXY') rating_text = rating_img['alt'] if rating_img else 'No rating available' # Extract and clean review count review_count_elem = card.find('span', class_='IiChw') review_count = review_count_elem.get_text(strip=True).split()[0] if review_count_elem else '0' # Format output exactly as requested print(f"{cafe_name}, {rating_text}, {review_count} reviews.") # Run the script target_url = 'https://www.tripadvisor.com.au/Restaurants-g255068-c8-Brisbane_Brisbane_Region_Queensland.html' tripadvisor_html = get_html(target_url) extract_cafe_info(tripadvisor_html)
Key Improvements Explained
- User-Agent Header: TripAdvisor blocks requests without a proper user-agent, so this prevents 403 Forbidden errors.
- Single HTML Parsing: We parse the HTML once instead of three times, making the script faster and more efficient.
- Linked Data Extraction: We grab each cafe's entire card first, so all info (name, rating, reviews) comes from the same entry—no mismatched data.
- Error Handling: Added checks for missing elements (like cafes with no ratings) to prevent the script from crashing.
- Cleaned Text: Used
get_text(strip=True)to remove extra whitespace, and formatted output to match your requested style perfectly.
Quick Tips for Future Scraping
- TripAdvisor’s HTML structure can change over time. If the script stops working, use your browser’s DevTools to inspect the page and update the class names.
- For larger scraping tasks, consider using the
requestslibrary instead ofurllib(it’s more user-friendly) and add small delays between requests to avoid being blocked.
内容的提问来源于stack exchange,提问作者noob123
相关产品推荐
相关产品推荐

