如何检验两个Probit模型中对应自变量系数的显著性差异?
Hey there! Let's break down exactly how to test whether the coefficients for your common predictors differ significantly across your two probit models. Since you're less familiar with Z-tests and related methods, I'll keep this practical and code-focused.
The Core Idea
We want to test the null hypothesis: the coefficient for a predictor (e.g., CyberOA) is the same in both models. To do this reliably (since both models use the same dataset), the best approach is to fit a single "joint" model that explicitly estimates the difference between coefficients across your two model specifications.
Step-by-Step Implementation in R
Here's how to do this using tidyverse and base R tools:
1. Prepare a Stacked Dataset
First, we'll create a marker variable to distinguish observations used in each model, then stack the datasets together:
# Load required packages (if you haven't already) library(dplyr) # Create two copies of your data, with a marker for each model df_model2 <- df27 %>% mutate(model_group = 1) # 1 = Model 2 (includes CyberCommittee) df_model2.1 <- df27 %>% mutate(model_group = 0, # 0 = Model 2.1 (no CyberCommittee) CyberCommittee = 0) # Fill with 0 to avoid missing values in the stacked data # Stack the two datasets stacked_data <- bind_rows(df_model2, df_model2.1)
2. Fit the Joint Probit Model
Next, we'll fit a probit model that includes:
- All predictors from your original models
- The
model_groupmarker - Interaction terms between each common predictor and
model_group
The interaction term coefficients will directly tell us the difference between coefficients across models, and their p-values will test if that difference is statistically significant:
# Fit the joint model joint_model <- glm( DataBreach ~ model_group + CyberCommittee + CyberOA*model_group + CyberJobs*model_group + CyberEducation*model_group + CyberAchievements*model_group, data = stacked_data, family = binomial(link = "probit") ) # View the results summary(joint_model)
How to Interpret the Results
- For each interaction term (e.g.,
CyberOA:model_group), the coefficient equals:(Coefficient in Model 2) - (Coefficient in Model 2.1) - The p-value associated with the interaction term tests whether this difference is statistically significant. A p-value < 0.05 means we reject the null hypothesis (i.e., the coefficients differ significantly between models).
Alternative: Manual Z-Test (Less Reliable)
If you prefer not to use the joint model approach, you can calculate the Z-statistic manually for each predictor:
# Extract coefficients and standard errors from both models coef_model2 <- coef(probit2) se_model2 <- summary(probit2)$coefficients[, "Std. Error"] coef_model2.1 <- coef(probit2.1) se_model2.1 <- summary(probit2.1)$coefficients[, "Std. Error"] # Calculate Z-statistic for each common predictor common_preds <- c("CyberOA", "CyberJobs", "CyberEducation", "CyberAchievements") for(pred in common_preds) { diff_coef <- coef_model2[pred] - coef_model2.1[pred] diff_se <- sqrt(se_model2[pred]^2 + se_model2.1[pred]^2) # Note: This assumes no covariance between estimates, which isn't strictly true for the same dataset z_stat <- diff_coef / diff_se p_val <- 2 * pnorm(-abs(z_stat)) # Two-tailed test cat(sprintf("%s: Z = %.3f, p-value = %.4f\n", pred, z_stat, p_val)) }
⚠️ Important: This manual approach ignores the covariance between coefficient estimates from the two models (since they use the same data), so the standard error calculation is approximate. The joint model method is more accurate.
Applying This to Your Results
Looking at your coefficient values:
- CyberAchievements: Difference = -0.026 - (-0.037) = 0.011
- CyberEducation: Difference = -0.040 - (-0.057) = 0.017
- CyberJobs: Difference = -0.025 - (-0.038) = 0.013
- CyberOA: Difference = -0.021 - (-0.029) = 0.008
Running the joint model will tell you if these differences are statistically meaningful, not just numerically different.
内容的提问来源于stack exchange,提问作者Robbert

