如何计算Aalen Additive模型系数的p值?lifelines包技术咨询
Hey there, I’ve worked with Cameron Davidson-Pilon’s lifelines package quite a bit, so let’s break down how to compute p-values for your regression coefficients—especially for the trickier Aalen Additive Model.
For standard models (like Cox Proportional Hazards)
While lifelines doesn’t surface p-values directly in some default outputs, you can derive them using the model’s built-in covariance matrix and basic statistical tests. Here’s a step-by-step example:
from lifelines import CoxPHFitter import scipy.stats as stats import pandas as pd import numpy as np # Assume you've already fitted your Cox model cph = CoxPHFitter() cph.fit(your_dataframe, duration_col="time_to_event", event_col="event_occurred") # Extract key model outputs coefficients = cph.params_ covariance_matrix = cph.covariance_matrix_ # Calculate standard errors (square root of diagonal elements in covariance matrix) standard_errors = np.sqrt(np.diag(covariance_matrix)) # Compute z-scores and two-tailed p-values z_scores = coefficients / standard_errors p_values = 2 * stats.norm.sf(np.abs(z_scores)) # Package results into a readable DataFrame results_df = pd.DataFrame({ "Coefficient": coefficients, "Standard Error": standard_errors, "Z-Score": z_scores, "p-value": p_values }) print(results_df)
This works because coefficients follow an approximate normal distribution under the null hypothesis, so we can use the standard normal distribution to calculate p-values.
For the Aalen Additive Model
The Aalen model is trickier because coefficients are time-varying—they change at each event time. You have two main options here: calculating p-values for coefficients at individual time points, or running a global test to check if coefficients are non-zero overall.
1. Time-point specific p-values
Here’s how to compute p-values for each coefficient at every event time:
from lifelines import AalenAdditiveFitter # Fit the Aalen model aaf = AalenAdditiveFitter() aaf.fit(your_dataframe, duration_col="time_to_event", event_col="event_occurred") # Extract time-varying coefficients and covariance matrices beta_time_varying = aaf.beta_survival_ covariance_time_varying = aaf.covariance_matrix_ # Calculate standard errors and p-values for each time point standard_errors = pd.DataFrame( [np.sqrt(np.diag(covariance_time_varying.loc[t])) for t in covariance_time_varying.index], index=covariance_time_varying.index, columns=beta_time_varying.columns ) z_scores_time_varying = beta_time_varying / standard_errors p_values_time_varying = 2 * stats.norm.sf(np.abs(z_scores_time_varying)) # View the first few time points' p-values print(p_values_time_varying.head())
This gives you a snapshot of how each covariate’s significance changes over time.
2. Global significance test
If you want to test whether a covariate has a non-zero effect overall, you can use a Wald test. Here’s a quick implementation:
# Use the final time point's coefficients and covariance matrix (or average coefficients if preferred) final_coefficients = beta_time_varying.iloc[-1].values final_covariance = covariance_time_varying.iloc[-1].values # Compute Wald statistic (follows a chi-squared distribution) wald_statistic = final_coefficients.T @ np.linalg.inv(final_covariance) @ final_coefficients global_p_value = stats.chi2.sf(wald_statistic, df=len(final_coefficients)) print(f"Global Wald test p-value (all covariates): {global_p_value}")
A few quick notes to keep in mind:
- Time-varying p-values need careful interpretation— a covariate might be significant early in the timeline but not later, or vice versa.
- Watch out for multicollinearity in your data: if covariates are highly correlated, the covariance matrix might be singular, leading to unstable p-value calculations. You can check this with variance inflation factor (VIF) tests.
内容的提问来源于stack exchange,提问作者Peadar Coyle

