You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按优先级合并DataFrame相邻列:每4列合并为1列

问题:按4列一组压缩DataFrame并按优先级取值

需求:将列数少于100的DataFrame按4列一组合并压缩,每组合并为1列,取值遵循以下优先级规则:

  • 最高优先级:组内任意一列值为1,合并结果为1
  • 次优先级:组内无1但存在0(其余可为NA),合并结果为0
  • 最低优先级:组内所有列均为NA,合并结果为NA

示例

输入数据

ID c1 c2 c3 c4 c5 c6 c7 c8
row1  1  0  0  0  1  0  0 NA
row2 NA NA NA  0 NA NA NA NA

输出结果

ID c1 c2
row1  1  1
row2  0 NA

可复现输入数据(R代码)

df <- structure(list(
  ID = c("row1", "row2"),
  c1 = c(1, NA), c2 = c(0, NA), c3 = c(0, NA), c4 = c(0, 0),
  c5 = c(1, NA), c6 = c(0, NA), c7 = c(0, NA), c8 = c(NA, NA)
), class = "data.frame", row.names = c(NA, -2L))

解决方案(R代码)

# 分离ID列与数据列
data_cols <- df[, -which(names(df) == "ID"), drop = FALSE]

# 生成4列一组的分组索引
group_indices <- ceiling(seq_along(data_cols) / 4)

# 按组处理每行数据,应用优先级规则
compressed_data <- t(apply(data_cols, 1, function(row) {
  tapply(row, group_indices, function(group) {
    if (any(group == 1, na.rm = TRUE)) {
      1
    } else if (any(group == 0, na.rm = TRUE)) {
      0
    } else {
      NA
    }
  })
}))

# 重命名压缩后的列
colnames(compressed_data) <- paste0("c", seq_len(ncol(compressed_data)))

# 合并ID列与压缩数据
result_df <- cbind(df[, "ID", drop = FALSE], as.data.frame(compressed_data))

# 输出结果
print(result_df)

代码说明

  1. 先拆分ID列和业务数据列,避免分组逻辑干扰ID列
  2. 通过ceiling(seq_along(data_cols)/4)生成每列对应的组编号,实现4列一组的自动分组
  3. 对每行数据按组依次校验优先级:先判断是否存在1,再判断是否存在0,最后返回NA
  4. 转置处理结果并重命名列,最后合并ID列得到最终的压缩DataFrame

内容的提问来源于stack exchange,提问作者Beardedant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 06:50:27