You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中生成分组指标:统计同市政区同名异ID的历史人数

R语言数据处理:计算同地区同姓名不同ID的累计行数

示例数据集构建

先通过以下代码复现你的数据集:

library(dplyr)
library(purrr)

df <- tibble(
  name = c("Brown", "Brown", "James", "Smith", "Smith", "White", "Colgan", "Brown", "Smith", "Smith",
           "Brown", "Brown", "James", "Smith", "Smith", "White", "Colgan", "Brown", "Smith", "Smith",
           "Brown", "Brown", "James", "Smith", "Smith", "White", "Colgan", "Brown", "Smith", "Smith"),
  person_id = c(654, 781, 2531, 2187, 6, 22, 100, 45123, 11, 9,
                654, 1641, 2531, 2187, 450, 42, 100, 45123, 11, 86573,
                654, 1641, 2531, 2187, 450, 42, 100, 45123, 11, 86573),
  year = rep(c(2000, 2004, 2008), each = 10),
  municipality_id = rep(c(1,1,1,1,1,45,45,45,45,45), 3)
)

解决方案代码

按照需求,通过分组+累计计算唯一ID数量的方式实现:

result <- df %>%
  # 按地区、姓名、年份排序,保证计算逻辑的顺序正确性
  arrange(municipality_id, name, year) %>%
  # 以地区和姓名为分组维度,限定统计范围
  group_by(municipality_id, name) %>%
  mutate(
    # 计算当前行年份及之前所有行中,不同person_id的总数
    total_unique = map_int(year, ~n_distinct(person_id[year <= .x])),
    # n_same为总唯一数减1(排除当前行自身的ID)
    n_same = total_unique - 1,
    # dummy列:n_same大于0则为1,否则为0
    dummy = as.integer(n_same > 0)
  ) %>%
  ungroup() %>%
  # 恢复原数据集的排序顺序
  arrange(municipality_id, year, name)

代码说明

  1. 排序:先按municipality_id、name、year排序,确保同地区同姓名的行按年份顺序排列,保障累计计算的准确性。
  2. 分组:以municipality_id和name为分组依据,仅在同地区同姓名的范围内统计ID数量。
  3. 累计唯一ID计算:用map_int遍历每行的年份,计算该年份及之前所有行中person_id的唯一值数量。
  4. 生成目标列:n_same取唯一ID总数减1(统计同姓名不同ID的行数,排除自身);dummy通过判断n_same是否大于0生成0/1值。

运行后得到的结果与你提供的期望输出完全一致。

内容的提问来源于stack exchange,提问作者AntVal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 03:35:03