You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R DataFrame按分组统计唯一值并计算月份间差异百分比

R实现按产品类别计算月度标签唯一值差异百分比

我们可以通过dplyr包的分组汇总能力快速实现需求,代码可直接运行,计算逻辑完全匹配规则:

# 加载依赖包
library(dplyr)

# 构造示例数据
df <- data.frame(
  stringsAsFactors = FALSE,
  month = c("jan","jan","jan","jan",
            "jan","feb","feb","feb","feb"),
  category = c("TB", "GT", "TB", "YT", "GT", "TB", "GT", "TB", "YT"),
  tag_number = c(101L, 101L, 223L, 223L, 223L, 345L, 345L, 655L, 223L)
)

# 计算差异百分比
result <- df %>%
  group_by(category) %>%
  summarise(
    # 提取1月、2月对应类别的唯一标签
    jan_tags = list(unique(tag_number[month == "jan"])),
    feb_tags = list(unique(tag_number[month == "feb"])),
    # 计算两月总唯一标签数
    total_unique = n_distinct(c(unlist(jan_tags), unlist(feb_tags))),
    # 计算两月标签交集数量
    intersect_cnt = length(intersect(unlist(jan_tags), unlist(feb_tags))),
    # 按规则计算差异百分比并格式化为带%的字符串
    pct_diff = paste0(round((total_unique - intersect_cnt)/total_unique * 100), "%")
  ) %>%
  select(category, pct_diff)

# 输出结果
print(result)

如果不想引入第三方包,也可以用基础R实现:

result <- do.call(rbind, lapply(split(df, df$category), function(sub_df) {
  jan_tags <- unique(sub_df$tag_number[sub_df$month == "jan"])
  feb_tags <- unique(sub_df$tag_number[sub_df$month == "feb"])
  total_unique <- length(unique(c(jan_tags, feb_tags)))
  intersect_cnt <- length(intersect(jan_tags, feb_tags))
  pct_diff <- paste0(round((total_unique - intersect_cnt)/total_unique * 100), "%")
  data.frame(category = unique(sub_df$category), pct_diff = pct_diff, row.names = NULL)
}))
rownames(result) <- NULL
print(result)

两种方案运行后输出结果均和预期完全一致:

category pct_diff
  <chr>    <chr>   
1 GT       100%    
2 TB       100%    
3 YT       0%  

内容的提问来源于stack exchange,提问作者Forge

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 16:24:03