Python Beautiful Soup提取元素问题:无法关联评分与对应维度
Hey there! Let's figure out how to reliably map those ratings (Cuisine, Service, Quality) to their respective scores. Based on common restaurant rating HTML structures and the issues you mentioned, here's a step-by-step solution to fix both the "no results" and "unmatched ratings" problems.
First, Let's Assume a Typical Target HTML Structure
Since you mentioned attaching the HTML, I'll use a common pattern for rating sections (adjust selectors later if your actual HTML differs):
<div class="restaurant-ratings"> <div class="rating-group"> <span class="rating-label">Cuisine</span> <span class="rating-value">4.5</span> </div> <div class="rating-group"> <span class="rating-label">Service</span> <span class="rating-value">4.0</span> </div> <div class="rating-group"> <span class="rating-label">Quality</span> <span class="rating-value">4.5</span> </div> </div>
Core Solution: Bind Labels to Ratings via Parent Elements
The key mistake in your second code was probably extracting all labels and all scores separately (e.g., find_all('span', class_='rating-label') and find_all('span', class_='rating-value')), which can break if the page structure changes. Instead, we'll:
- Grab each individual rating's parent container
- Extract the label and score within that container to guarantee they're paired correctly
Example Code (Python + BeautifulSoup)
from bs4 import BeautifulSoup # Replace with your actual HTML content (or fetch it via requests/Selenium) html = """ <div class="restaurant-ratings"> <div class="rating-group"> <span class="rating-label">Cuisine</span> <span class="rating-value">4.5</span> </div> <div class="rating-group"> <span class="rating-label">Service</span> <span class="rating-value">4.0</span> </div> <div class="rating-group"> <span class="rating-label">Quality</span> <span class="rating-value">4.5</span> </div> </div> """ # Parse the HTML soup = BeautifulSoup(html, 'html.parser') # Step 1: Get all parent containers for each rating pair rating_containers = soup.find_all('div', class_='rating-group') # Step 2: Map labels to scores final_ratings = {} for container in rating_containers: # Extract label (adjust selector to match your HTML's label element) label = container.find('span', class_='rating-label').get_text(strip=True) # Extract score (adjust selector to match your HTML's score element) score = container.find('span', class_='rating-value').get_text(strip=True) final_ratings[label] = score # Output the matched results print(final_ratings) # Expected output: {'Cuisine': '4.5', 'Service': '4.0', 'Quality': '4.5'}
Fixing the "No Results" Issue
If your first code returned nothing, check these common causes:
- Wrong selectors: Double-check class names, tag types, or nested structures (e.g., maybe the rating labels are in
<p>instead of<span>). - Dynamic content: If the ratings load via JavaScript (e.g., after page scroll or AJAX), use tools like Selenium to fetch the fully rendered HTML:
from selenium import webdriver from selenium.webdriver.common.by import By from bs4 import BeautifulSoup driver = webdriver.Chrome() driver.get("your-target-url-here") driver.implicitly_wait(5) # Wait for dynamic content to load html = driver.page_source soup = BeautifulSoup(html, 'html.parser') # Use the same rating extraction code as above driver.quit()
内容的提问来源于stack exchange,提问作者yellow days

