Statsmodels GLM分析AB测试泊松回归时误报完全分离的解决方法
我尝试用泊松回归分析AB测试结果,先通过以下代码模拟数据:
import numpy as np import pandas as pd import statsmodels.api as sm def make_ab_test_data(n_samples, pi_c, lift): pi_t = pi_c * lift y_t = np.random.binomial(n=n_samples, p=pi_t) y_c = np.random.binomial(n=n_samples, p=pi_c) y = [y_t, y_c] X = sm.add_constant([1, 0]) N = np.ones_like(y) * n_samples return y, X, N y, X, N = make_ab_test_data(10_000, 0.01, 1.1)
使用Statsmodels的GLM函数分析时,运行sm.GLM(y, X, family = sm.families.Poisson(), offset = np.log(N)).fit()会抛出PerfectSeparationError: Perfect separation detected, results not available错误,但相同数据在R中运行同类模型可正常得到结果:
# Example dataset N <- rep(10000, 2) y <- c(113, 81) txt <-c(1, 0) glm(y~txt + offset(log(N)), family = poisson()) |> summary()
R的输出结果:
Call:
glm(formula = y ~ txt + offset(log(N)), family = poisson())Deviance Residuals:
[1] 0 0Coefficients:
Estimate Std. Error z value Pr(>|z|)
(Intercept) -4.8159 0.1111 -43.343 <2e-16 ***
txt 0.3329 0.1456 2.287 0.0222 *Signif. codes: 0 ‘’ 0.001 ‘’ 0.01 ‘’ 0.05 ‘.’ 0.1 ‘ ’ 1
(Dispersion parameter for poisson family taken to be 1)
Null deviance: 5.3026e+00 on 1 degrees of freedomResidual deviance: -1.3545e-14 on 0 degrees of freedom
AIC: 16.801Number of Fisher Scoring iterations: 2
解决方法
Statsmodels默认会检测完全分离并抛出错误,但这类仅两组的AB测试数据(每组对应1个聚合观测点)不存在真正的分离问题,只是模型完全拟合了数据(残差为0)。只需在拟合时禁用相关检查即可:
model = sm.GLM(y, X, family=sm.families.Poisson(), offset=np.log(N)) results = model.fit(check_pearson=False, check_singular=False) print(results.summary())
执行后会得到和R一致的拟合结果。另外,也可以将数据展开为个体级观测(每个用户对应一行,事件列标记0或1)再拟合,但这种方式对大样本会消耗更多内存,不如直接调整拟合参数高效。
内容的提问来源于stack exchange,提问作者Demetri Pananos

