R语言:为数据框添加颜色频率列与最频繁颜色列
需求:扩展数据框的统计列功能
这是此前问题的延伸,现需在已有实现基础上添加更多统计列。
现有数据框df
ID <- c(1,1,1,1,1,1,1,2,2,2,2,2) color <- c("red","red","red","blue","green","green","blue", "yellow","yellow","red","blue","green") df <- data.frame(ID,color)
输出:
ID color 1 1 red 2 1 red 3 1 red 4 1 blue 5 1 green 6 1 green 7 1 blue 8 2 yellow 9 2 yellow 10 2 red 11 2 blue 12 2 green
已实现的功能
已成功创建n_distinct_color列,用于展示每个ID对应的不同颜色数量:
df %>% group_by(ID) %>% distinct(color, .keep_all = T) %>% mutate(n_distinct_color = n(), .after = ID) %>% ungroup()
输出:
# A tibble: 7 × 3 ID n_distinct_color color <dbl> <int> <chr> 1 1 3 red 2 1 3 blue 3 1 3 green 4 2 4 yellow 5 2 4 red 6 2 4 blue 7 2 4 green
需要新增的列
需在上述结果基础上添加两个新列:
frequency_of_color:显示每个ID下各颜色的出现次数(如ID1中red出现3次,blue出现2次)most_frequent_color:显示每个ID的最频繁颜色(如ID1的最频繁颜色是red,ID2是yellow)
期望输出
ID n_distinct_color color frequency_of_color most_frequent_color <dbl> <int> <chr> <int> <chr> 1 1 3 red 3 red 2 1 3 blue 2 red 3 1 3 green 2 red 4 2 4 yellow 2 yellow 5 2 4 red 1 yellow 6 2 4 blue 1 yellow 7 2 4 green 1 yellow
多最频颜色的情况
若存在多个颜色频率相同的情况(如ID2中yellow和red频率相同),需明确数据框的输出形式。示例数据框df_new如下:
ID <- c(1,1,1,1,1,1,1,2,2,2,2,2,2) color <- c("red","red","red","blue","green","green","blue", "yellow","yellow","red","blue","green","red") df_new <- data.frame(ID,color)
输出:
ID color 1 1 red 2 1 red 3 1 red 4 1 blue 5 1 green 6 1 green 7 1 blue 8 2 yellow 9 2 yellow 10 2 red 11 2 blue 12 2 green 13 2 red
内容的提问来源于stack exchange,提问作者Bruh
相关产品推荐
相关产品推荐

