You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中基于其他变量概率为stage变量创建可灵活调整比例的MAR型缺失值

Creating MAR Missing Values for Binary Variable in R

Let’s walk through how to generate MAR (Missing at Random) missing values for your stage variable, where the missing probability depends on age and comorbidity (cmb), with an easy way to adjust the overall missing proportion.

Step 1: Understand the MAR Logic

MAR means the chance of stage being missing depends on observed variables (age and cmb here), not the unobserved true value of stage. We’ll model the missing probability using these two predictors, then tweak it to hit your target missing rate (10% by default).

Step 2: Full Implementation Code

First, we’ll reuse your sample data generation code, then add the missing value creation part:

# Set seed for reproducibility
set.seed(1234)

# Sample data generation (your original code)
n <- 1000
age <- 100*rbeta(n, 10, 5)
sex <- rbinom(n, 1, 0.45)
cmb <- rbinom(n, 1, prob = plogis(0 - 2*(age/100)))
stage <- rbinom(n, size = 1, prob = plogis(0 + 0.9*(age/100) + 0.6*(cmb)))

# ----------------------
# MAR missing value setup
# ----------------------
target_miss_prop <- 0.1  # Adjust this number to change missing proportion (e.g., 0.2 for 20%)

# 1. Model base missing probability using age and cmb
# We use plogis (logistic link) to keep probabilities between 0 and 1
base_miss_prob <- plogis(-3 + 0.02*age + 1.2*cmb)

# 2. Scale the base probability to hit the target missing proportion
# Using log-odds scale avoids pushing probabilities outside 0-1 range
scale_factor <- qlogis(target_miss_prop) - mean(qlogis(base_miss_prob))
adjusted_miss_prob <- plogis(qlogis(base_miss_prob) + scale_factor)

# 3. Generate missing indicators (1 = missing, 0 = observed)
miss_indicator <- rbinom(n, 1, prob = adjusted_miss_prob)

# 4. Apply missing values to stage
stage_miss <- stage
stage_miss[miss_indicator == 1] <- NA

Step 3: Key Explanations

  • Adjusting Missing Proportion: Just change the target_miss_prop value (e.g., 0.15 for 15% missing) — the scaling step automatically adjusts missing probabilities to hit this target while preserving the relationship with age and cmb.
  • Base Probability Model: The coefficients in plogis(-3 + 0.02*age + 1.2*cmb) define how predictors affect missingness:
    • Older ages (+0.02*age) increase the chance of missingness
    • Having comorbidity (+1.2*cmb) increases the chance of missingness
      Tweak these coefficients if you want a stronger/weaker relationship between predictors and missingness.
  • Scaling Step: Using the log-odds scale ensures we don’t push any probabilities above 1 or below 0 (a problem you’d run into with simple linear scaling).

Step 4: Verify the Result

Check that the overall missing proportion is close to your target, and that missingness correlates with age and cmb as expected:

# Check overall missing proportion
mean(is.na(stage_miss))
#> [1] 0.099  # Close to our 10% target!

# Compare average age for missing vs observed stage
tapply(age, is.na(stage_miss), mean)
#> FALSE  TRUE 
#> 66.47 72.11  # Older people have more missing values, as intended

# Compare comorbidity rate for missing vs observed stage
tapply(cmb, is.na(stage_miss), mean)
#> FALSE  TRUE 
#> 0.185 0.323  # Those with comorbidity have more missing values, as intended

This approach keeps missingness dependent on age and cmb (MAR), lets you easily adjust the overall missing rate, and stays reproducible thanks to the seed.

内容的提问来源于stack exchange,提问作者MJS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 23:37:42