如何使用dplyr将分组输出结果整理为指定结构的命名列表
需求说明
需要将如下tibble格式的df数据转换为指定格式的命名列表:列表元素名对应df的name列取值,每个元素是以value为键、freq为值的命名向量。
测试数据
df <- tibble( name = c("ageThenCode", "ageThenCode", "ageThenCode", "ageThenCode", "ageThenCode", "ageThenCode", "eduCode", "eduCode", "eduCode", "regioCat", "regioCat", "regioCat", "teltipCode", "teltipCode", "teltipCode"), value = factor(c("1", "2", "3", "4", "5", "6", "1", "2", "3", "1", "2", "3", "1", "2", "3")), freq = c(0.042, 0.222, 0.208, 0.222, 0.149, 0.157, 0.155, 0.591, 0.254, 0.373, 0.34, 0.287, 0.271, 0.524, 0.205) )
预期输出格式
list( ageThenCode = c("1"=0.042, "2"=0.222, "3"=0.208, "4"=0.222, "5"=0.149, "6"=0.157), eduCode = c("1"=.155, "2"=.591, "3"=.254), regionCode = c("1"=.373, "2"=.34, "3"=.287), teltipCode = c("1"=.271, "2"=.524,"3"=.205) )
注:预期输出中eduCode的数值为示例近似值,实际以df中freq列为准;df中regioCat对应预期输出的regionCode,可以提前修改列值对齐命名要求。
错误尝试代码
df %>% group_by(name) %>% summarise(cat = list(value=freq) ) %>% pivot_wider(names_from = name, values_from = cat)
正确实现方法
方法1:tidyverse 实现(适配原有写法的修改)
核心问题是没有将value设置为freq向量的命名,调整如下:
library(tidyverse) result <- df %>% # 可选:如果需要把regioCat改成regionCode,先加一步重命名 # mutate(name = recode(name, "regioCat" = "regionCode")) %>% group_by(name) %>% # 每组生成一个以value为名字、freq为值的命名向量,存为列表元素 summarise(vec = list(set_names(freq, as.character(value)))) %>% # 直接将两列的tibble转换为命名列表,第一列为列表名,第二列为列表元素 deframe()
运行后打印result即可得到符合要求的输出,可通过result$ageThenCode验证单组结果。
方法2:基础R实现(不需要加载额外包)
result <- split(df, df$name) |> lapply(function(x) setNames(x$freq, as.character(x$value))) # 可选重命名regioCat # names(result)[names(result) == "regioCat"] <- "regionCode"
内容的提问来源于stack exchange,提问作者SunWuKung
相关产品推荐
相关产品推荐

