如何从Sephora网站的aria-label标签中提取产品评分?
Hey there! I see you're stuck on grabbing the product rating from Sephora's pages — let's get that sorted out. The main issue with your current code is that soup.find('a', {'aria-label'}) only fetches the HTML element itself, not the actual rating value stored in the aria-label attribute. Plus, Sephora's rating elements have a specific structure we need to target properly.
Let's Break Down the Problem
Sephora's product rating lives in an <a> tag with an aria-label that looks like this: aria-label="4.7 stars, 1000+ reviews". Your original code wasn't extracting the numeric rating from this attribute, which is why the extraction failed.
Updated Code with Fixes
Here's your revised code with two key improvements:
- Properly extracting the rating from the
aria-labelattribute - Adding a user-agent header to avoid being blocked by Sephora's anti-scraping measures
import requests from bs4 import BeautifulSoup final_products = [] # Empty list to append product dictionaries before DataFrame conversion headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } for current_url in link_list: # Simplified the loop (no need for index range) try: resource = requests.get(current_url, headers=headers) resource.raise_for_status() # Raise error for bad HTTP responses (e.g., 403, 404) current_data = resource.text soup = BeautifulSoup(current_data, 'html.parser') product = {} # Extract product name name_elem = soup.find('span', {'class': 'css-0'}) product['Name'] = name_elem.text.strip() if name_elem else None # Extract price price_elem = soup.find('div', {'class': 'css-slwsq8'}) product['Price'] = price_elem.text.strip() if price_elem else None # Extract number of reviews review_elems = soup.find_all('span', {'class': 'css-2rg6q7'}) product['Num_of_reviews'] = review_elems[0].text.strip() if review_elems else None # Extract number of likes likes_elem = soup.find('span', {'data-at': 'product_love_count'}) product['Num_of_Likes'] = likes_elem.text.strip() if likes_elem else None # Extract rating (fixed part!) rating_elem = soup.find('a', {'aria-label': lambda attr: attr and 'stars' in attr.lower()}) if rating_elem: aria_label = rating_elem['aria-label'] # Split the label to get the numeric rating (e.g., "4.7" from "4.7 stars...") product['Rating'] = aria_label.split()[0] else: product['Rating'] = None # Mark as None if no rating found final_products.append(product) except requests.exceptions.RequestException as e: print(f"Error fetching {current_url}: {e}") except Exception as e: print(f"Error parsing {current_url}: {e}")
Key Changes Explained
- User-Agent Header: Sephora blocks requests without a proper user-agent, so adding this mimics a real browser visit to avoid being flagged as a crawler.
- Targeted Rating Extraction: We use a lambda function in
find()to locate the<a>tag with anaria-labelcontaining "stars", then extract the first part of that attribute (the numeric rating). - Graceful Error Handling: Instead of a broad
except: pass, we now catch specific request errors and general parsing errors, so you can debug issues without silently skipping pages. - None Fallbacks: For each field, we check if the element exists before accessing
.text, preventing AttributeErrors if a field is missing on a product page.
Quick Notes
- Sephora's HTML classes might change over time, so if this breaks later, inspect the product page again to update the class names or attributes.
- Consider adding small delays between requests (
time.sleep(2)) to avoid overwhelming Sephora's servers and getting blocked.
内容的提问来源于stack exchange,提问作者DJ-coding

