R语言处理亚马逊销售排名字符串:提取指定排名整数
在R中处理亚马逊商品销售排名字符串的解决方案
准备工作
先确保安装并加载stringr包(tidyverse生态的一部分,字符串处理语法更直观):
# 未安装则先执行 # install.packages("stringr") library(stringr)
步骤1:清理并拆分排名条目
先处理原始数据中的换行符(示例数据含换行,实际场景也可能遇到),再按分号拆分每个商品的多个排名条目:
# 构造示例数据集 df <- data.frame( rank = c(">#316,475 in Cell Phones & Accessories ;>#1,908 in Cell Phones & Accessories > Cell Phone Accessories > Headphones > On-Ear Headphones;#3,410 in Cell Phones & Accessories > Cell Phone Accessories > Headphones > Over-Ear Headphones;#52,046 in Electronics > Portable Audio & Video > MP3 & MP4 Player Accessories", "...") ) # 清理换行符,将每个商品的排名拆分为条目列表 df$rank_items <- str_split(str_replace_all(df$rank, "\\n", " "), ";")
步骤2:筛选目标分类的排名条目
假设要筛选包含On-Ear Headphones的排名条目,用str_detect匹配关键词:
# 定义目标分类关键词 target_category <- "On-Ear Headphones" # 从每个商品的排名条目中筛选匹配项 df$target_rank_entry <- sapply(df$rank_items, function(items) { items[str_detect(items, target_category)] })
步骤3:提取排名整数数值
从筛选出的条目中提取带逗号的数字,移除逗号后转换为整数:
# 提取并转换排名数值 df$rank_number <- sapply(df$target_rank_entry, function(entry) { if(length(entry) == 0) { NA # 无匹配条目时返回NA } else { # 提取数字部分,去掉#和逗号,转整数 as.integer(str_remove_all(str_extract(entry, "#\\d+,\\d+"), "#|,")) } })
结果示例
运行后df$rank_number列会得到目标分类的排名数值,比如示例中第一个商品的结果为1908。
扩展说明
- 更换目标分类只需修改
target_category的值 - 若排名数字无逗号,可简化提取逻辑:
as.integer(str_remove(str_extract(entry, "#\\d+"), "#")) - 批量处理多分类可结合
purrr包实现循环筛选
内容的提问来源于stack exchange,提问作者Jesse Brefeld
相关产品推荐
相关产品推荐

