R语言CSS选择器爬取网页:原价与折扣价提取问题求助
解决R爬虫提取折扣价的问题
问题分析
你当前提取折扣价时用的选择器.b-product-tile-price-item会同时匹配到带--line-through类的原价元素,导致数据混乱;另外品牌选择器过于宽泛,容易抓取到无关内容。
修正后的代码
library(rvest) library(dplyr) # 读取目标页面 url <- "https://www.snipes.ch/fr/c/shoes?srule=Standard&prefn1=isSale&prefv1=true&openCategory=true&sz=48" html <- read_html(url) # 定位每个商品的容器 html_products <- html %>% html_elements("div.b-product-tile-info-container") # 提取商品名称 product_names <- html_products %>% html_element(".b-product-tile-link") %>% html_text2() # 自动清理换行、多余空格 # 提取商品品牌(用精准类选择器定位品牌元素) product_brand <- html_products %>% html_element(".b-product-tile-brand") %>% html_text2() # 提取原价 product_op <- html_products %>% html_element(".b-product-tile-price-item--line-through") %>% html_text2() # 提取折扣价:排除带划线的原价元素 product_dp <- html_products %>% html_element(".b-product-tile-price-item:not(.b-product-tile-price-item--line-through)") %>% html_text2() # 生成DataFrame product_df <- tibble( 品牌 = product_brand, 商品名称 = product_names, 原价 = product_op, 折扣价 = product_dp ) print(product_df)
关键修正点
- 折扣价选择器:使用
:not()伪类排除带划线的原价元素,确保只抓取折扣价。 - 文本清理:用
html_text2()替代html_text(),自动处理换行、多余空格,无需额外gsub操作。 - 品牌选择器:改用精准的
.b-product-tile-brand类,避免抓取到其他无关的<span>元素。
内容的提问来源于stack exchange,提问作者sikimimi
相关产品推荐
相关产品推荐

