如何将列表中DataFrame的列值设为对应DataFrame的名称
问题:将列表中DataFrame的指定列值设为列表元素名称并移除该列
之前可以用以下dplyr代码将列表中各DataFrame的名称存入名为id的列:
library(dplyr) bind_rows(list(A = df1, B = df2), .id = 'id')
现在需要执行反向操作:当列表中每个DataFrame的指定列(如示例中的Species列)存储着对应DataFrame的目标名称时,将该列值设为列表中对应DataFrame的名称,并移除该列。
示例输入my_list
[[1]] # A tibble: 2 x 5 Sepal.Length Sepal.Width Petal.Length Petal.Width Species <dbl> <dbl> <dbl> <dbl> <chr> 1 5.1 3.5 1.4 0.2 new_setoas 2 4.9 3 1.4 0.2 new_setoas [[2]] # A tibble: 2 x 5 Sepal.Length Sepal.Width Petal.Length Petal.Width Species <dbl> <dbl> <dbl> <dbl> <chr> 1 6.3 3.3 6 2.5 new_virginica 2 5.8 2.7 5.1 1.9 new_virginica [[3]] # A tibble: 2 x 5 Sepal.Length Sepal.Width Petal.Length Petal.Width Species <dbl> <dbl> <dbl> <dbl> <chr> 1 7 3.2 4.7 1.4 versicolor 2 6.4 3.2 4.5 1.5 versicolor my_list <- structure(list(structure(list(Sepal.Length = c(5.1, 4.9), Sepal.Width = c(3.5, 3), Petal.Length = c(1.4, 1.4), Petal.Width = c(0.2, 0.2), Species = c("new_setoas", "new_setoas")), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -2L)), structure(list(Sepal.Length = c(6.3, 5.8), Sepal.Width = c(3.3, 2.7), Petal.Length = c(6, 5.1), Petal.Width = c(2.5, 1.9), Species = c("new_virginica", "new_virginica")), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -2L)), structure(list(Sepal.Length = c(7, 6.4), Sepal.Width = c(3.2, 3.2), Petal.Length = c(4.7, 4.5), Petal.Width = c(1.4, 1.5), Species = c("versicolor", "versicolor")), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -2L))), ptype = structure(list( Sepal.Length = numeric(0), Sepal.Width = numeric(0), Petal.Length = numeric(0), Petal.Width = numeric(0), Species = character(0)), class = c("tbl_df", "tbl", "data.frame"), row.names = integer(0)), class = c("vctrs_list_of", "vctrs_vctr", "list"))
期望输出
[[new_setoas]] # A tibble: 2 x 5 Sepal.Length Sepal.Width Petal.Length Petal.Width <dbl> <dbl> <dbl> <dbl> 1 5.1 3.5 1.4 0.2 2 4.9 3 1.4 0.2 [[new_virginica]] # A tibble: 2 x 5 Sepal.Length Sepal.Width Petal.Length Petal.Width <dbl> <dbl> <dbl> <dbl> 1 6.3 3.3 6 2.5 2 5.8 2.7 5.1 1.9 [[versicolor]] # A tibble: 2 x 5 Sepal.Length Sepal.Width Petal.Length Petal.Width <dbl> <dbl> <dbl> <dbl> 1 7 3.2 4.7 1.4 2 6.4 3.2 4.5 1.5
解决方案
方法1:使用purrr + dplyr
借助tidyverse工具链,遍历列表提取目标名称、移除指定列后重命名列表:
library(purrr) library(dplyr) # 指定要处理的列名 target_col <- "Species" # 执行操作 result <- my_list %>% # 移除每个DataFrame的指定列 map(~ .x %>% select(-all_of(target_col))) %>% # 用指定列的第一个值作为列表元素的新名称 set_names(map_chr(my_list, ~ first(.x[[target_col]])))
方法2:基础R实现
不依赖额外包,用基础R函数完成相同操作:
target_col <- "Species" # 提取每个DataFrame的目标名称(假设该列值全一致) new_names <- sapply(my_list, function(df) df[[target_col]][1]) # 移除指定列并设置列表名称 result <- lapply(my_list, function(df) df[, !names(df) %in% target_col]) names(result) <- new_names
注:两种方法均默认每个DataFrame的指定列值完全一致,取第一个值作为列表元素的新名称。
内容的提问来源于stack exchange,提问作者TarJae
相关产品推荐
相关产品推荐

