如何用Python BeautifulSoup从仅含class的图标元素提取星级评分?
问题分析与解决
原代码的核心错误
- 容器元素未匹配:如果
soup.findAll('div', attrs={'class':'_6uN7R'})返回空列表,循环根本不会执行,直接导致ratings为空。这大概率是因为目标页面的class名已更新,或者页面评分区域是JS动态加载的,静态爬取(BeautifulSoup)无法获取。 - 星星元素获取错误:
f.find('i')只会返回第一个<i>标签,无法通过[4]索引取到第5个星星,必须用find_all('i')获取所有星星元素。 - class判断逻辑错误:
hasattr(rating, "_9-ogB fqfC4")是检查元素对象是否有名为_9-ogB fqfC4的属性,这完全不符合需求,正确的做法是检查元素的class列表中是否包含目标类名。
修正方案
情况1:静态页面可获取评分元素
如果确认页面是静态渲染(可通过查看网页源代码找到评分标签),用以下代码:
ratings = [] # 先确保外层容器的class正确,若不对需重新检查页面元素 for item in soup.find_all('div', class_='_6uN7R'): rating_container = item.find('div', class_='mdmmT _32vUv') if not rating_container: # 没有评分的商品,可添加默认值如0 ratings.append(0.0) continue # 获取所有星星元素 stars = rating_container.find_all('i') full_stars = 0 half_stars = 0 for star in stars: star_classes = star.get('class', []) if 'Dy1nx' in star_classes: full_stars +=1 elif 'fqfC4' in star_classes: half_stars +=1 # 计算最终评分 total_rating = full_stars + half_stars * 0.5 ratings.append(total_rating) print(ratings)
情况2:页面动态加载评分
Lazada多数商品数据是JS动态加载的,静态爬取无法获取到_6uN7R这类容器元素,此时需要用Selenium模拟浏览器渲染:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC ratings = [] driver = webdriver.Chrome() driver.get('https://www.lazada.com.my/catalog/?q=live+plants&_keyori=ss&from=input&spm=..search.go.') # 等待商品容器加载完成 try: WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CLASS_NAME, '_6uN7R')) ) items = driver.find_elements(By.CLASS_NAME, '_6uN7R') for item in items: try: rating_container = item.find_element(By.CLASS_NAME, 'mdmmT._32vUv') stars = rating_container.find_elements(By.TAG_NAME, 'i') full_stars = 0 half_stars = 0 for star in stars: star_classes = star.get_attribute('class').split() if 'Dy1nx' in star_classes: full_stars +=1 elif 'fqfC4' in star_classes: half_stars +=1 total_rating = full_stars + half_stars *0.5 ratings.append(total_rating) except: # 无评分商品添加默认值 ratings.append(0.0) finally: driver.quit() print(ratings)
额外提示
- 爬取电商平台需遵守网站的
robots.txt规则,避免频繁请求导致IP被封禁。 - 若class名再次变化,需重新检查页面元素,更新对应的选择器。
内容的提问来源于stack exchange,提问作者edger02
相关产品推荐
相关产品推荐

