You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于条件填充纵向数据集中缺失值的技术咨询

Fill Missing Values in Longitudinal Dataset by ID and Year

Got it, let's work through filling those missing marstat values in your longitudinal dataset. First, let's recap your data structure to confirm we're on the same page—you have repeated observations per individual (id) and year, with some marstat entries missing, and we want to fill those NAs using the non-missing value from the same id and year group (since a person's marital status should be consistent within a single year).

First, let's recreate your dataset to test the solutions:

id <- c(rep("1", 5), rep("2", 5), rep("3", 5))
year <- c(1999, 1999, 2000, 2001, 2001, 1999, 2000, 2001, 2001, 2001, 1999, 2000, 2001, 2002, 2003)
marstat <- c("married", NA, "married", "married", "divorced", "single", "single", "single", NA, NA, "married", NA, "married", "divorced", "divorced")
df <- data.frame(id , year , marstat)

Solution 1: Using dplyr (Tidyverse Approach)

This is a clean, readable way to handle grouping and filling:

library(dplyr)

# Group by id and year, then fill NAs with the non-missing value from the group
df_filled <- df %>%
  group_by(id, year) %>%
  mutate(
    marstat = case_when(
      !is.na(marstat) ~ marstat,
      # If there's at least one non-NA value in the group, use it
      length(na.omit(marstat)) > 0 ~ first(na.omit(marstat)),
      # If the whole group is NA, leave it as NA
      TRUE ~ marstat
    )
  ) %>%
  ungroup()

# Check the result
print(df_filled)

Solution 2: Base R Approach (No External Packages)

If you prefer not to load tidyverse packages, use ave() to handle grouping and filling:

df_filled_base <- df

# Use ave() to apply a filling function per id-year group
df_filled_base$marstat <- with(df_filled_base,
  ave(marstat, id, year, FUN = function(x) {
    # Get non-missing values in the group
    non_missing <- na.omit(x)
    # If there are non-missing values, repeat them for the whole group; else keep NAs
    if (length(non_missing) > 0) {
      rep(non_missing, length(x))
    } else {
      x
    }
  })
)

# Check the result
print(df_filled_base)

Key Notes

  • Both solutions handle cases where an id-year group has all NAs (those will remain NA since there's no valid value to fill with).
  • We assume that within a single id and year, all non-missing marstat values are identical (which makes sense for marital status in a given year). If you have conflicting non-missing values in the same group, you'll need to add logic to resolve that (e.g., take the most frequent value, but that's unlikely here).

内容的提问来源于stack exchange,提问作者T.P.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:44:55