多水平模型中的非劣效性分析:随机临床试验技术问询
Great question—combining multilevel models (MLMs) with noninferiority testing is such a smart move here. Instead of running separate TOST at each time point (which wastes the repeated measures structure and inflates Type I error), you can integrate noninferiority logic directly into your MLM framework to get more robust, efficient results. Let’s break this down for both your continuous and binary outcomes:
Core Idea
Instead of testing each time point in isolation, we’ll embed your noninferiority hypothesis into the MLM’s parameter inference. This leverages the correlation between repeated measures, reduces estimation variance, and lets you analyze both overall and time-specific noninferiority in one unified model. The key is to compare group differences against your pre-specified noninferiority margin (δ)—a clinically meaningful threshold that defines "acceptable" difference between your two therapies.
1. Continuous Outcome (Pain Intensity)
Start with a standard linear multilevel model that models within-person changes over time and between-person group differences:
# Example R formula for lme4 library(lme4) model_continuous <- lmer(Pain ~ Group * Time + (1 + Time | Subject), data = your_data)
Where:
Group: Binary indicator (1 = experimental, 0 = control)Time: Coded as 0 (baseline) to 4 (final follow-up)(1 + Time | Subject): Random intercept and slope to account for individual variation in baseline pain and pain change over time
Noninferiority Inference Steps
- Define your δ: Let’s say clinical experts agree a pain difference ≤ 1 point (on a 0-10 scale) is noninferior—so δ = 1.
- Focus on key group difference parameters:
- For time-specific noninferiority: Calculate the predicted pain difference between groups at each time point using marginal means. For time
t, this difference isΔ_t = γ_01 + γ_11*t(whereγ_01is the baseline group difference,γ_11is the group-time interaction). - For overall noninferiority: Compute the average group difference across all time points, or test if the overall group effect (averaged over time) stays within δ.
- For time-specific noninferiority: Calculate the predicted pain difference between groups at each time point using marginal means. For time
- Confidence Interval (CI) Approach (equivalent to TOST for MLMs):
- For one-sided noninferiority (experimental ≤ control + δ), check if the 95% upper one-sided CI of
Δ_t(or the average Δ) is ≤ δ. If yes, you can conclude noninferiority. - If you need two-sided noninferiority (i.e., experimental is no worse than control by δ, and no better than control by δ), use a two-sided CI and verify it falls within [-δ, δ].
- For one-sided noninferiority (experimental ≤ control + δ), check if the 95% upper one-sided CI of
Implementation Tip
Use the emmeans package to easily compute marginal means and CIs for group differences at each time point:
library(emmeans) time_point_comparisons <- emmeans(model_continuous, pairwise ~ Group | Time) # Extract the contrasts and their CIs time_point_comparisons$contrasts
Then just compare the upper CI bound to your δ.
2. Binary Outcome (Adverse Events)
For binary outcomes, use a generalized linear multilevel model (GLMM) with a logit or probit link:
model_binary <- glmer(AdverseEvent ~ Group * Time + (1 + Time | Subject), data = your_data, family = binomial)
Noninferiority Inference Steps
- Translate δ to the model scale: Noninferiority for binary outcomes is usually defined via risk difference (RD) or odds ratio (OR). For example:
- If you accept a ≤20% higher adverse event risk in the experimental group: δ_RD = 0.2
- Or an OR ≤1.25 (experimental odds are ≤25% higher): δ_OR = 1.25 (log(δ_OR) ≈ 0.223 for the logit scale)
- Compute marginal effects: Unlike conditional means (which depend on random effects), we need marginal means (population-averaged effects) to compare group outcomes at the population level.
- CI-based testing:
- For RD: Calculate the marginal risk difference at each time point, then check if its 95% upper one-sided CI ≤ δ_RD.
- For OR: Compute the OR (exponentiate the logit-scale group difference) and verify its 95% upper one-sided CI ≤ δ_OR.
Implementation Tip
Use emmeans with the type = "response" argument to get marginal probabilities/RDs on the original scale:
binary_comparisons <- emmeans(model_binary, pairwise ~ Group | Time, type = "response") binary_comparisons$contrasts
Key Considerations
- δ must be clinically justified: Don’t pick a margin based on statistics alone—work with clinicians to define what counts as "noninferior" for your outcome.
- Model fit matters: Make sure your random structure (e.g., random slopes for time) matches your data. If patients’ pain trajectories vary a lot, skipping random slopes will bias your results.
- Multiple testing: If you’re testing noninferiority at 5 time points, consider adjusting your CI confidence level (e.g., use 99% CIs instead of 95% to account for 5 comparisons) or focus on an overall group-time interaction effect instead of individual time points.
- Why this is better than separate TOST: MLMs account for within-person correlation, so your estimates are more precise. You also avoid inflating Type I error from running 5 separate tests, and can easily model how noninferiority changes over time.
内容的提问来源于stack exchange,提问作者Peter Flom

