使用Tidyverse统计各颜色对应的id组数及重要颜色关联id数
问题与解法
主问题
输入数据
library(tibble) df <- tibble( id = c("101", "101", "101", "102", "102", "103"), color = c("Blue", "Blue", "Red", "Red", "Green", "Green") )
数据预览:
# A tibble: 6 × 2 id color <chr> <chr> 1 101 Blue 2 101 Blue 3 101 Red 4 102 Red 5 102 Green 6 103 Green
需求
统计每个color对应的拥有该颜色的唯一id数量,预期输出:
# A tibble: 3 × 2 color number_of_ids_having_color <chr> <chr> 1 Blue 1 2 Red 2 3 Green 2
Tidyverse解法
library(dplyr) df %>% group_by(color) %>% summarise(number_of_ids_having_color = n_distinct(id)) %>% mutate(number_of_ids_having_color = as.character(number_of_ids_having_color))
附加问题
输入数据
df2 <- tibble( id = c("101", "101", "101", "102", "102", "103"), color = c("Blue", "Blue", "Red", "Red", "Green", "Green"), most_important_color = c("F", "F", "T", "T", "F", "T") )
数据预览:
# A tibble: 6 × 3 id color most_important_color <chr> <chr> <chr> 1 101 Blue F 2 101 Blue F 3 101 Red T 4 102 Red T 5 102 Green F 6 103 Green T
需求
在同一管道中同时统计:
- 每个
color对应的拥有该颜色的唯一id数 - 该颜色作为对应
id的most_important_color(值为"T")的唯一id数
预期输出:
# A tibble: 3 × 3 color_group number_of_ids_having_color number_of_ids_having_color_as_most_important <chr> <chr> <chr> 1 Blue 1 0 2 Red 2 2 3 Green 2 1
Tidyverse解法
df2 %>% group_by(color) %>% summarise( number_of_ids_having_color = n_distinct(id), number_of_ids_having_color_as_most_important = n_distinct(id[most_important_color == "T"]) ) %>% rename(color_group = color) %>% mutate(across(c(number_of_ids_having_color, number_of_ids_having_color_as_most_important), as.character))
内容的提问来源于stack exchange,提问作者daltoncito5034
相关产品推荐
相关产品推荐

