You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言CSS选择器爬取网页:原价与折扣价提取问题求助

解决R爬虫提取折扣价的问题

问题分析

你当前提取折扣价时用的选择器.b-product-tile-price-item会同时匹配到带--line-through类的原价元素,导致数据混乱;另外品牌选择器过于宽泛,容易抓取到无关内容。

修正后的代码

library(rvest)
library(dplyr)

# 读取目标页面
url <- "https://www.snipes.ch/fr/c/shoes?srule=Standard&prefn1=isSale&prefv1=true&openCategory=true&sz=48"
html <- read_html(url)

# 定位每个商品的容器
html_products <- html %>% html_elements("div.b-product-tile-info-container")

# 提取商品名称
product_names <- html_products %>% 
  html_element(".b-product-tile-link") %>% 
  html_text2()  # 自动清理换行、多余空格

# 提取商品品牌(用精准类选择器定位品牌元素)
product_brand <- html_products %>% 
  html_element(".b-product-tile-brand") %>% 
  html_text2()

# 提取原价
product_op <- html_products %>% 
  html_element(".b-product-tile-price-item--line-through") %>% 
  html_text2()

# 提取折扣价:排除带划线的原价元素
product_dp <- html_products %>% 
  html_element(".b-product-tile-price-item:not(.b-product-tile-price-item--line-through)") %>% 
  html_text2()

# 生成DataFrame
product_df <- tibble(
  品牌 = product_brand,
  商品名称 = product_names,
  原价 = product_op,
  折扣价 = product_dp
)

print(product_df)

关键修正点

  • 折扣价选择器:使用:not()伪类排除带划线的原价元素,确保只抓取折扣价。
  • 文本清理:用html_text2()替代html_text(),自动处理换行、多余空格,无需额外gsub操作。
  • 品牌选择器:改用精准的.b-product-tile-brand类,避免抓取到其他无关的<span>元素。

内容的提问来源于stack exchange,提问作者sikimimi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 07:27:12