按优先级合并DataFrame相邻列:每4列合并为1列
问题:按4列一组压缩DataFrame并按优先级取值
需求:将列数少于100的DataFrame按4列一组合并压缩,每组合并为1列,取值遵循以下优先级规则:
- 最高优先级:组内任意一列值为
1,合并结果为1 - 次优先级:组内无
1但存在0(其余可为NA),合并结果为0 - 最低优先级:组内所有列均为NA,合并结果为NA
示例
输入数据
ID c1 c2 c3 c4 c5 c6 c7 c8 row1 1 0 0 0 1 0 0 NA row2 NA NA NA 0 NA NA NA NA
输出结果
ID c1 c2 row1 1 1 row2 0 NA
可复现输入数据(R代码)
df <- structure(list( ID = c("row1", "row2"), c1 = c(1, NA), c2 = c(0, NA), c3 = c(0, NA), c4 = c(0, 0), c5 = c(1, NA), c6 = c(0, NA), c7 = c(0, NA), c8 = c(NA, NA) ), class = "data.frame", row.names = c(NA, -2L))
解决方案(R代码)
# 分离ID列与数据列 data_cols <- df[, -which(names(df) == "ID"), drop = FALSE] # 生成4列一组的分组索引 group_indices <- ceiling(seq_along(data_cols) / 4) # 按组处理每行数据,应用优先级规则 compressed_data <- t(apply(data_cols, 1, function(row) { tapply(row, group_indices, function(group) { if (any(group == 1, na.rm = TRUE)) { 1 } else if (any(group == 0, na.rm = TRUE)) { 0 } else { NA } }) })) # 重命名压缩后的列 colnames(compressed_data) <- paste0("c", seq_len(ncol(compressed_data))) # 合并ID列与压缩数据 result_df <- cbind(df[, "ID", drop = FALSE], as.data.frame(compressed_data)) # 输出结果 print(result_df)
代码说明
- 先拆分ID列和业务数据列,避免分组逻辑干扰ID列
- 通过
ceiling(seq_along(data_cols)/4)生成每列对应的组编号,实现4列一组的自动分组 - 对每行数据按组依次校验优先级:先判断是否存在
1,再判断是否存在0,最后返回NA - 转置处理结果并重命名列,最后合并ID列得到最终的压缩DataFrame
内容的提问来源于stack exchange,提问作者Beardedant
相关产品推荐
相关产品推荐

