You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中通过记录扩展实现记录关联:统一实体最小ID

解决方案

原始数据:

library(tibble)
df <- tibble(id = c(1, 2, 2, 3, 4),
             x = c(123, 123, 125, 125, 200))

这个问题本质是连通分量匹配:把id和x看作图中的节点,每个id-x对是一条边,属于同一连通分量的所有id都应该被替换为该分量中最小的id。可以用igraph包高效处理这种关联问题:

library(tidyverse)
library(igraph)

# 1. 构建边列表:每个id和对应的x建立连接
edges <- df %>%
  select(id, x) %>%
  mutate(x = paste0("x_", x))  # 给x加前缀避免和id数值冲突

# 2. 创建图并提取连通分量
graph <- graph_from_data_frame(edges, directed = FALSE)
components <- components(graph)$membership

# 3. 提取每个连通分量中的最小id
component_min_id <- enframe(components, name = "node", value = "group") %>%
  filter(str_detect(node, "^\\d+$")) %>%  # 筛选出id节点
  mutate(node = as.integer(node)) %>%
  group_by(group) %>%
  summarise(min_id = min(node)) %>%
  right_join(enframe(components, name = "node", value = "group"), by = "group") %>%
  select(node, min_id)

# 4. 映射回原始数据,替换id
result <- df %>%
  mutate(id_str = as.character(id),
         x_str = paste0("x_", x)) %>%
  # 匹配id对应的min_id
  left_join(component_min_id %>% filter(str_detect(node, "^\\d+$")), by = c("id_str" = "node")) %>%
  # 匹配x对应的min_id
  left_join(component_min_id %>% filter(str_detect(node, "^x_")), by = c("x_str" = "node")) %>%
  # 取非空的min_id
  mutate(new_id = coalesce(min_id.x, min_id.y)) %>%
  select(id = new_id, x)

print(result)

运行结果:

# A tibble: 5 × 2
     id     x
  <int> <dbl>
1     1   123
2     1   123
3     1   125
4     1   125
5     4   200

说明

  • 给x添加前缀x_是为了避免x的数值和id重复导致节点混淆;
  • 连通分量处理会自动把所有通过x关联的id归为一组:比如id=1通过x=123关联id=2,id=2又通过x=125关联id=3,三者属于同一分量,取最小的id=1;
  • 如果没有id=1的记录,id=2和id=3会因x=125关联,归为同一组并取最小id=2,完全符合需求。

内容的提问来源于stack exchange,提问作者sparklink

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 14:47:20