You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中为每个ID提取首诊年份并填充至所有行

解决方案:按ID统一填充首次确诊年份

你的核心需求是按ID分组,将该ID最早确诊的年份填充到所有对应行,而非逐行计算。下面提供两种简洁的实现方式:

方法1:长格式转换+合并(直观易理解)

先将宽格式数据转为长格式,筛选出确诊记录后按ID取最小年份,再合并回原数据:

library(dplyr)
library(tidyr)

# 原始数据集
df <- data.frame(
  id = rep(1:5, each = 5),
  physical_1987 = c(1, 0, 0, 0, 0),
  physical_1988 = c(0, 1, 0, 0, 0),
  physical_1989 = c(0, 0, 1, 0, 0),
  physical_1990 = c(1, 0, 0, 0, 0),
  physical_1991 = c(0, 0, 0, 0, 1)
)

# 1. 转长格式,提取年份并筛选确诊记录
df_first_year <- df %>%
  pivot_longer(
    cols = starts_with("physical_"), 
    names_to = "year_str", 
    values_to = "diagnosed"
  ) %>%
  mutate(year = as.integer(sub("physical_", "", year_str))) %>%
  filter(diagnosed == 1) %>%
  group_by(id) %>%
  summarize(first_physical = min(year, na.rm = TRUE)) %>%
  ungroup()

# 2. 合并回原数据,无确诊记录的ID自动填充NA
df <- df %>% left_join(df_first_year, by = "id")

# 查看结果
df

方法2:分组直接计算(无需转格式)

利用dplyr的分组功能,在每组内直接计算最早确诊年份:

library(dplyr)

df <- data.frame(
  id = rep(1:5, each = 5),
  physical_1987 = c(1, 0, 0, 0, 0),
  physical_1988 = c(0, 1, 0, 0, 0),
  physical_1989 = c(0, 0, 1, 0, 0),
  physical_1990 = c(1, 0, 0, 0, 0),
  physical_1991 = c(0, 0, 0, 0, 1)
)

# 获取所有年份列名
year_cols <- starts_with("physical_", vars = names(df))

df <- df %>%
  group_by(id) %>%
  mutate(
    first_physical = {
      # 遍历年份列,提取该ID有确诊记录的年份
      positive_years <- sapply(year_cols, function(col) {
        if (any(.data[[col]] == 1)) {
          as.integer(sub("physical_", "", col))
        } else {
          NA_integer_
        }
      })
      # 取最小的确诊年份,无则返回NA
      min(positive_years, na.rm = TRUE)
    }
  ) %>%
  ungroup()

# 查看结果
df

两种方法都能实现需求:比如id=1的所有行first_physical都会被填充为1987,id=5的所有行填充为1991,无确诊记录的ID则为NA。

内容的提问来源于stack exchange,提问作者user25809482

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 06:43:16