基于全样本生成倾向评分(PS)的亚样本调整分析技术咨询
Got it, let's walk through the key technical implementation points you need to nail when applying full-sample-generated propensity scores (PS) to your 40% subsample—while keeping your approach concise and enabling direct comparisons with unadjusted results. Here's what to focus on:
1. Lock in consistent PS generation from the full sample
First and foremost, never refit the PS model using only the subsample. Your PS must be estimated using the entire study population's covariates (all the variables you included in the original full-sample PS model, e.g., demographics, baseline characteristics). This ensures the PS reflects the exposure assignment probability across your entire cohort, not just the subset you're analyzing later.
For example, in R, this would look like:
# Fit PS model on the FULL dataset full_ps_model <- glm(exposure ~ age + gender + baseline_score + comorbidities, data = full_cohort, family = binomial()) # Predict PS values for your 40% subsample using the full-sample model subsample$ps <- predict(full_ps_model, newdata = subsample, type = "response")
- Critical note: If there are missing covariates in the subsample, use the same missing data handling strategy you applied to the full sample (e.g., multiple imputation models trained on the full cohort). Inconsistent missing data treatment will skew your PS predictions.
2. Align model specifications across analyses
To make valid comparisons between adjusted/unadjusted results (and between subsample/full-sample results), keep your regression model structures identical:
- For the subsample's PS-adjusted model: Use the exact same formula as your full-sample PS covariate adjustment model (e.g.,
outcome ~ exposure + ps—or if you used a polynomial/PS quintiles in the full sample, replicate that here). - For the subsample's unadjusted model: Match the full-sample unadjusted model (e.g.,
outcome ~ exposure), no extra covariates added. - If your 40% subsample is a targeted subgroup (not random), explicitly document its definition (e.g., "patients with follow-up >12 months") so readers understand the context of your subset analysis.
3. Verify balance and run quick sensitivity checks
Even with full-sample PS, you need to confirm it still balances covariates in the subsample:
- Plot PS density curves for exposed vs. unexposed groups in the subsample to check overlap. If one group's PS range is drastically narrower, your balance might be compromised.
- Calculate standardized mean differences (SMD) for all covariates before/after PS adjustment in the subsample. Aim for SMD < 0.1 for all variables—this confirms the full-sample PS is effectively reducing selection bias in the subset.
- Compare effect estimates across three scenarios: subsample unadjusted, subsample PS-adjusted, and full-sample PS-adjusted. Note any discrepancies (e.g., wider confidence intervals in the subsample due to smaller size) and explain potential drivers.
4. Be transparent in reporting
When writing up your results, avoid ambiguity:
- Clearly state that the PS used in the subsample analysis was estimated from the entire study cohort, not the subsample itself.
- Present side-by-side effect estimates (e.g., odds ratios, hazard ratios) and confidence intervals for all three analysis types to highlight how adjustment and subsampling impact your findings.
- In the discussion, acknowledge the tradeoffs: using full-sample PS gives you a more stable estimate of exposure assignment probability than fitting a model on the smaller subsample, but if the subsample has fundamentally different covariate distributions than the full cohort, this approach may have limitations (though for an open queue, this is less likely).
内容的提问来源于stack exchange,提问作者Vincent

