如何在R中为数据框创建统计物品总频次的新列
在R中实现物品全局频次统计并生成新列
问题分析
你的原代码使用str_count()得到的是每行中物品的出现次数,而需求是获取整个列中物品的总出现频次,并将该频次作为固定值填充到对应新列的每一行。
解决方案
可以通过以下步骤实现需求:
- 从
items列中提取所有唯一的物品名称 - 对每个物品,计算其在整个列中的总出现次数
- 将这些频次作为新列添加到原数据框中
代码实现
library(dplyr) library(stringr) # 定义原始数据框 items <- structure(list(items = c("Apple", "Apple, Pear", "Apple, Pear, Banana")), row.names = c(NA, -3L), class = "data.frame") # 提取所有唯一物品 all_unique_items <- items %>% pull(items) %>% str_split(", ") %>% unlist() %>% unique() # 生成包含全局频次的新列 result_df <- items %>% mutate(across(all_unique_items, ~ sum(str_detect(items, .x)))) # 查看结果 result_df
代码说明
str_split(", ") %>% unlist():将每行的物品字符串拆分为单个物品,再转换为一维向量unique():获取所有不重复的物品名称across(all_unique_items, ~ sum(str_detect(items, .x))):遍历每个唯一物品,用str_detect()判断每行是否包含该物品,sum()统计总次数,最终将每个物品的总频次作为新列添加到数据框中
结果验证
运行代码后得到的数据框与预期结构一致:
structure(list(items = c("Apple", "Apple, Pear", "Apple, Pear, Banana"), Apple = c(3, 3, 3), Pear = c(2, 2, 2), Banana = c(1, 1, 1)), row.names = c(NA, -3L), class = "data.frame")
内容的提问来源于stack exchange,提问作者hachiko
相关产品推荐
相关产品推荐

