You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr在R中比较两列值并生成匹配标记变量(含NA处理)

使用case_when实现列值匹配判断生成新列

直接用dplyr的case_when就能实现需求,关键是把NA的判断放在最前面(case_when按顺序匹配条件),再依次处理匹配和不匹配的情况:

library(dplyr)

# 你的原始示例数据
df = data.frame(id  = c(1, 2, 3),
                Test = c(3, 0, 1),
                More  = c(4, 0, 0))

# 生成Match列
df <- df %>%
  mutate(Match = case_when(
    # 任意一列是NA则返回NA
    is.na(Test) | is.na(More) ~ NA_real_,
    # 两列值相等返回1
    Test == More ~ 1L,
    # 其他情况(非NA且不相等)返回0
    TRUE ~ 0L
  ))

print(df)

运行后输出:

id Test More Match
1  1    3    4     0
2  2    0    0     1
3  3    1    0     0

补充NA场景测试

如果数据里包含NA,这个逻辑同样适用,比如修改数据加入NA值:

df_with_na = data.frame(id  = c(1, 2, 3, 4),
                        Test = c(3, NA, 1, NA),
                        More  = c(4, 0, NA, NA))

df_with_na <- df_with_na %>%
  mutate(Match = case_when(
    is.na(Test) | is.na(More) ~ NA_real_,
    Test == More ~ 1L,
    TRUE ~ 0L
  ))

print(df_with_na)

输出:

id Test More Match
1  1    3    4     0
2  2   NA    0    NA
3  3    1   NA    NA
4  4   NA   NA    NA

说明

  • NA_real_用于保证列的数值类型一致,避免类型冲突;如果偏好整数型,也可以用NA_integer_,同时保持1和0为1L/0L(示例中已用整数型)。
  • case_when的条件顺序很重要,必须先判断NA的情况,否则NA会被后续的Test == More判断为不匹配,返回0,不符合需求。

内容的提问来源于stack exchange,提问作者whosurdata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 12:21:49