You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:基于另一数据集的值为数据集新增Owner列

使用dplyr为df2匹配最接近分数的Owner

示例数据构造

先还原你提供的两个数据集:

library(dplyr)

df1 <- tibble(
  Id = c("00001", "00008", "00011", "00004"),
  Owner = c("John", "Paul", "Alba", "Martha"),
  Remaining_points = c(18, 34, 52, 67)
)

df2 <- tibble(
  Id = c("00025", "00076", "00089", "00092"),
  Points = c(17, 35, 51, 68)
)

解决方案代码

通过交叉连接计算差值,筛选最接近的Owner:

df_result <- df2 %>%
  # 交叉连接df1,让每个df2的行与所有df1的行配对
  cross_join(df1) %>%
  # 计算Points与Remaining_points的绝对差值
  mutate(diff = abs(Points - Remaining_points)) %>%
  # 按df2的Id分组
  group_by(Id) %>%
  # 筛选出差值最小的行(with_ties=FALSE确保只取一个,若有多个最小差可调整)
  slice_min(order_by = diff, n = 1, with_ties = FALSE) %>%
  # 保留需要的列
  select(Id, Points, Owner) %>%
  ungroup()

print(df_result)

输出结果

运行后会得到你期望的数据集:

# A tibble: 4 × 3
  Id    Points Owner 
  <chr>  <dbl> <chr> 
1 00025     17 John  
2 00076     35 Paul  
3 00089     51 Alba  
4 00092     68 Martha

补充说明

如果存在多个Owner与某个Points的差值相同的情况,可以调整slice_min的参数:

  • 若想保留所有最小差的记录,去掉with_ties=FALSE
  • 若需要自定义优先级(比如取名字首字母靠前的),可以在order_by里添加额外排序条件,比如order_by = c(diff, Owner)

内容的提问来源于stack exchange,提问作者kikusanchez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 08:25:23