如何将含标识列的data.frame按列分区生成指定结构的列表
我来帮你实现这个实用的分区函数foo,它既支持指定分区数量,也能通过自定义每部分的非标识列数来分割数据框,完全匹配你的需求!
函数实现
foo <- function(df, partitions = NULL, columns_per_element = NULL, identifier_col = 1) { # 参数校验:避免同时传入两种分区参数 if (!is.null(partitions) && !is.null(columns_per_element)) { stop("请仅传入 `partitions` 或 `columns_per_element` 其中一个参数哦") } # 分离标识列与非标识列(drop=FALSE 保证提取后仍是数据框格式) id_col <- df[, identifier_col, drop = FALSE] non_id_cols <- df[, -identifier_col, drop = FALSE] total_non_id <- ncol(non_id_cols) # 处理「自定义每部分列数」的场景 if (!is.null(columns_per_element)) { if (sum(columns_per_element) != total_non_id) { stop("`columns_per_element` 的总和必须等于非标识列的总数哦") } # 生成列分割的索引位置 split_points <- cumsum(c(1, columns_per_element)) split_points <- split_points[-length(split_points)] # 分割非标识列 split_non_id <- split(non_id_cols, cut(seq_len(total_non_id), breaks = split_points, labels = FALSE)) } # 处理「指定分区数」的场景:自动平均分配列数 else if (!is.null(partitions)) { if (partitions > total_non_id) { stop("分区数不能超过非标识列的数量哦") } # 计算每个分区应分配的列数(尽可能平均) base_cols <- floor(total_non_id / partitions) extra_cols <- total_non_id %% partitions col_counts <- rep(base_cols, partitions) col_counts[1:extra_cols] <- col_counts[1:extra_cols] + 1 # 复用自定义列数的逻辑,减少代码重复 return(foo(df, columns_per_element = col_counts, identifier_col = identifier_col)) } # 未传入有效参数的情况 else { stop("请传入 `partitions` 或 `columns_per_element` 参数哦") } # 将每个分区的非标识列与标识列合并,最终返回列表 lapply(split_non_id, function(part_cols) cbind(id_col, part_cols)) }
测试示例
先创建你的示例数据框(设置随机种子让结果可复现):
set.seed(123) df <- data.frame(id = 1:10, x1 = rnorm(10), x2 = rnorm(10), x3 = rnorm(10), x4 = rnorm(10))
场景1:指定分区数为3
result_partitions <- foo(df, partitions = 3) str(result_partitions)
输出结构和你预期的完全一致:
List of 3 $ :'data.frame': 10 obs. of 2 variables: ..$ id: int [1:10] 1 2 3 4 5 6 7 8 9 10 ..$ x1: num [1:10] -0.5605 -0.2302 1.5587 0.0705 0.1293 ... $ :'data.frame': 10 obs. of 2 variables: ..$ id: int [1:10] 1 2 3 4 5 6 7 8 9 10 ..$ x2: num [1:10] 1.715 0.461 -1.265 -0.687 -0.446 ... $ :'data.frame': 10 obs. of 3 variables: ..$ id: int [1:10] 1 2 3 4 5 6 7 8 9 10 ..$ x3: num [1:10] 1.224 0.359 0.401 0.111 -0.556 ... ..$ x4: num [1:10] 0.4008 0.1107 -0.5552 1.7869 0.4979 ...
场景2:自定义每部分列数c(1,1,2)
result_custom <- foo(df, columns_per_element = c(1,1,2)) str(result_custom)
结果和场景1完全相同,满足你的扩展需求。
额外小功能
函数还支持指定标识列的列名(而不是默认第一列),比如你的标识列叫user_id,可以这样调用:
foo(df, partitions=2, identifier_col="user_id")
内容的提问来源于stack exchange,提问作者socialscientist
相关产品推荐
相关产品推荐

