You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中从数据框列表中按规则筛选行?

问题说明

先创建两个数据框并组成列表:

df1 = data.frame(
  y1 = c(1, -2, 3),
  y2 = c(4, 5, 6),
  out = c("*", "*", "")
)

df2 = data.frame(
  y1 = c(7, -3, 9),
  y2 = c(1, 4, 6),
  out = c("", "*", "")
)

lis = list(df1, df2)

需要逐行对比列表中两个数据框的对应行(lis[[1]]第n行 vs lis[[2]]第n行),遵循以下规则筛选结果:

  • 若两行的out列均为*或均为空:选取y1绝对值更大的行,新增ind列标注其所属数据框的索引(1或2)
  • 若两行out列不同:直接选取out为*的行,新增ind列标注索引

期望输出:

y1 y2 out ind
 1  4   *    1
-3  4   *    2
 9  6        2
实现方案

方法一:基础R循环实现

# 初始化结果数据框
result = data.frame(y1 = numeric(), y2 = numeric(), out = character(), ind = integer(), stringsAsFactors = FALSE)

# 逐行遍历处理
for (i in 1:nrow(df1)) {
  row1 = df1[i, ]
  row2 = df2[i, ]
  
  if (row1$out == row2$out) {
    # 两行out一致,选y1绝对值大的
    if (abs(row1$y1) >= abs(row2$y1)) {
      selected_row = cbind(row1, ind = 1)
    } else {
      selected_row = cbind(row2, ind = 2)
    }
  } else {
    # 两行out不同,选out为*的行
    selected_row = if (row1$out == "*") cbind(row1, ind = 1) else cbind(row2, ind = 2)
  }
  
  result = rbind(result, selected_row)
}

# 查看结果
print(result)

运行输出:

y1 y2 out ind
1   1  4   *   1
2  -3  4   *   2
3   9  6       2

方法二:dplyr+tidyr 简洁实现

如果习惯使用tidyverse工具链,可以用更简洁的代码实现:

library(dplyr)
library(tidyr)

# 合并数据框并添加索引,按原行号分组
combined_data = bind_rows(lis, .id = "ind") %>%
  mutate(ind = as.integer(ind)) %>%
  group_by(row_number())

# 按规则筛选目标行
result = combined_data %>%
  filter(
    # 情况1:out相同,取y1绝对值最大的行
    (out == first(out) & abs(y1) == max(abs(y1))) |
    # 情况2:out不同,取out为*的行
    (out != first(out) & out == "*")
  ) %>%
  ungroup() %>%
  select(y1, y2, out, ind)

print(result)

运行后同样得到符合要求的结果。

内容的提问来源于stack exchange,提问作者bic ton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 09:45:29