如何在R语言中正确提取网页中五星百分比宽度形式的产品平均评分?
如何正确提取产品平均星级评分的宽度值?
我需要从页面抓取产品平均评分,该评分以五星金色填充的百分比宽度形式展示。使用以下R代码尝试提取宽度值时,得到了9个不同的style_attribute值,其中2个为NA,其余均不是目标值(当前示例为width: 91.6%),请问如何正确提取仅平均星级评分的宽度值?
page <- read_html("http://www.gonser.ch/13879") # extract the div element div_element <- html_nodes(page, ".feedback-stars-overlay-wrap") # Extract the "style" attribute from the element style_attribute <- html_attr(div_element, "style") # extract the width value width_value <- str_extract(style_attribute, "width: ([0-9.]+)%") # Convert to a numeric value width <- as.numeric(width_value)
解决方案
问题出在你用的CSS选择器.feedback-stars-overlay-wrap匹配了页面中所有星级元素(包括每个用户评价的单独星级),所以返回了一堆无关结果。要精准获取平均评分的宽度,必须缩小选择范围,定位到页面顶部的评分汇总区域对应的星级元素。
修正后的代码如下:
library(rvest) library(stringr) # 读取目标页面 page <- read_html("https://www.gonser.ch/13879") # 精准定位平均评分的星级填充元素(仅匹配评分汇总区域内的目标元素) avg_star_element <- html_nodes(page, ".feedback-summary .feedback-stars-overlay-wrap") # 提取style属性内容 style_attr <- html_attr(avg_star_element, "style") # 提取并转换宽度数值 width_percent <- style_attr %>% str_extract("width:\\s*([0-9.]+)%") %>% # 匹配包含宽度的字符串 str_remove("width:\\s*") %>% # 移除"width: "前缀 str_remove("%") %>% # 移除百分号 as.numeric() # 转换为数值类型 # 输出结果 print(width_percent)
关键说明
- 选择器优化:
.feedback-summary .feedback-stars-overlay-wrap仅匹配页面顶部评分汇总区域内的星级填充元素,避免抓取到用户评价的单个星级。 - 字符串清理:通过
str_remove剔除多余的字符串内容,最终得到纯数值的百分比值。 - 若后续页面结构变动,可通过浏览器开发者工具查看平均星级元素的父级容器class,调整选择器即可适配。
内容的提问来源于stack exchange,提问作者Pippo
相关产品推荐
相关产品推荐

