如何在R中从数据框按技术人员及品种统计type占比?
解决方案
我们可以用dplyr包快速实现这两个统计需求,先确认你的原始数据:
df <- data.frame(tech=c("Leonardo", "Leonardo", "Leonardo", "John", "John", "John", "Will", "Will", "Will", "Bob"), type=c("V", "P", "V", "V", "P", "V", "V", "P", "V", "V"), breed=c("A", "A", "A", "B", "B", "B", "C", "C", "A", "B"), central=c("J", "J", "K", "J", "K", "J", "K", "K", "K", "J"))
1. 各技术人员的type(V/P)占比
按tech分组后,统计每个type的出现次数,再计算占比:
library(dplyr) # 生成第一个数据框 df_type_ratio <- df %>% group_by(tech, type) %>% count() %>% group_by(tech) %>% mutate(ratio = n / sum(n) * 100) %>% ungroup() # 查看结果 df_type_ratio
运行后会得到每个技术人员对应V和P的占比,比如Leonardo的V占比为66.67%、P为33.33%,Bob仅有一种type,占比100%。
2. 各技术人员按品种细分后的type占比
按tech+breed组合分组,再统计type的占比:
# 生成第二个数据框 df_breed_type_ratio <- df %>% group_by(tech, breed, type) %>% count() %>% group_by(tech, breed) %>% mutate(ratio = n / sum(n) * 100) %>% ungroup() # 查看结果 df_breed_type_ratio
这个结果会按「技术人员-品种」的组合统计type占比,比如Will的品种C中V和P各占50%,品种A中仅V,占比100%。如果不需要保留计数列n,可以在mutate后添加select(-n)移除。
内容的提问来源于stack exchange,提问作者pete
相关产品推荐
相关产品推荐

