You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中按列分组统计两列匹配与不匹配的数量及百分比?

在R语言中按分组统计两列匹配情况的实现方案

可以使用dplyr包高效完成这个分组统计需求,步骤如下:

1. 构造示例数据

先还原你的输入数据:

df <- data.frame(
  col1 = c("A", "A", "A", "G", "G", "Y", "Y", "J", "J", "J", "K"),
  col2 = c("A", "A", "B", "G", "H", "Y", "Y", "P", "P", "J", "L")
)

2. 分组统计并格式化结果

使用dplyr的分组、汇总和格式化功能实现需求:

library(dplyr)

result <- df %>%
  # 按col1分组
  group_by(col1) %>%
  # 计算各项统计指标
  summarise(
    count_match = sum(col1 == col2),
    count_notmatch = n() - count_match,
    percent_match = count_match / n() * 100,
    percent_notmatch = 100 - percent_match
  ) %>%
  # 适配示例输出的百分比格式
  mutate(
    percent_match = case_when(
      percent_match == 0 ~ "0",
      TRUE ~ sprintf("%.2f%%", percent_match)
    ),
    percent_notmatch = case_when(
      percent_notmatch == 0 ~ "0",
      TRUE ~ sprintf("%.2f%%", percent_notmatch)
    )
  )

# 打印结果(隐藏行号)
print(result, row.names = FALSE)

3. 输出结果

运行代码后会得到与你期望一致的统计结果:

col1 count_match count_notmatch percent_match percent_notmatch
    A           2              1        66.66%           33.33%
    G           1              1         50.00%            50.00%
    Y           2              0        100.00%                0
    J           1              2        33.33%           66.66%
    K           0              1             0           100.00%

代码说明

  • group_by(col1):指定按col1的取值分组计算
  • sum(col1 == col2):利用逻辑值的数值特性(TRUE=1,FALSE=0)统计匹配行数
  • n():返回每组总行数,用于计算不匹配数和百分比
  • sprintf("%.2f%%", x):将数值格式化为保留两位小数的百分比字符串,case_when处理0值的特殊格式需求

内容的提问来源于stack exchange,提问作者Rich

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 00:45:25