如何在R中基于其他变量概率为stage变量创建可灵活调整比例的MAR型缺失值
Let’s walk through how to generate MAR (Missing at Random) missing values for your stage variable, where the missing probability depends on age and comorbidity (cmb), with an easy way to adjust the overall missing proportion.
Step 1: Understand the MAR Logic
MAR means the chance of stage being missing depends on observed variables (age and cmb here), not the unobserved true value of stage. We’ll model the missing probability using these two predictors, then tweak it to hit your target missing rate (10% by default).
Step 2: Full Implementation Code
First, we’ll reuse your sample data generation code, then add the missing value creation part:
# Set seed for reproducibility set.seed(1234) # Sample data generation (your original code) n <- 1000 age <- 100*rbeta(n, 10, 5) sex <- rbinom(n, 1, 0.45) cmb <- rbinom(n, 1, prob = plogis(0 - 2*(age/100))) stage <- rbinom(n, size = 1, prob = plogis(0 + 0.9*(age/100) + 0.6*(cmb))) # ---------------------- # MAR missing value setup # ---------------------- target_miss_prop <- 0.1 # Adjust this number to change missing proportion (e.g., 0.2 for 20%) # 1. Model base missing probability using age and cmb # We use plogis (logistic link) to keep probabilities between 0 and 1 base_miss_prob <- plogis(-3 + 0.02*age + 1.2*cmb) # 2. Scale the base probability to hit the target missing proportion # Using log-odds scale avoids pushing probabilities outside 0-1 range scale_factor <- qlogis(target_miss_prop) - mean(qlogis(base_miss_prob)) adjusted_miss_prob <- plogis(qlogis(base_miss_prob) + scale_factor) # 3. Generate missing indicators (1 = missing, 0 = observed) miss_indicator <- rbinom(n, 1, prob = adjusted_miss_prob) # 4. Apply missing values to stage stage_miss <- stage stage_miss[miss_indicator == 1] <- NA
Step 3: Key Explanations
- Adjusting Missing Proportion: Just change the
target_miss_propvalue (e.g.,0.15for 15% missing) — the scaling step automatically adjusts missing probabilities to hit this target while preserving the relationship withageandcmb. - Base Probability Model: The coefficients in
plogis(-3 + 0.02*age + 1.2*cmb)define how predictors affect missingness:- Older ages (
+0.02*age) increase the chance of missingness - Having comorbidity (
+1.2*cmb) increases the chance of missingness
Tweak these coefficients if you want a stronger/weaker relationship between predictors and missingness.
- Older ages (
- Scaling Step: Using the log-odds scale ensures we don’t push any probabilities above 1 or below 0 (a problem you’d run into with simple linear scaling).
Step 4: Verify the Result
Check that the overall missing proportion is close to your target, and that missingness correlates with age and cmb as expected:
# Check overall missing proportion mean(is.na(stage_miss)) #> [1] 0.099 # Close to our 10% target! # Compare average age for missing vs observed stage tapply(age, is.na(stage_miss), mean) #> FALSE TRUE #> 66.47 72.11 # Older people have more missing values, as intended # Compare comorbidity rate for missing vs observed stage tapply(cmb, is.na(stage_miss), mean) #> FALSE TRUE #> 0.185 0.323 # Those with comorbidity have more missing values, as intended
This approach keeps missingness dependent on age and cmb (MAR), lets you easily adjust the overall missing rate, and stays reproducible thanks to the seed.
内容的提问来源于stack exchange,提问作者MJS

