使用statsmodels.formula.api进行多元Logit回归报错求助
解决statsmodels多元Logit回归
ValueError: endog has evaluated to an array with multiple columns问题 问题原因
statsmodels的mnlogit公式接口会自动将字符串类型的因变量转换为多列虚拟变量(one-hot编码),但mnlogit要求因变量是单列整数型分类标识(代表每个样本对应的类别索引),这一矛盾导致了报错。
修复步骤
1. 将字符串因变量转换为整数编码
用pandas的factorize()方法把字符串类别转换成从0开始的整数编码,同时保留原类别映射:
import pandas as pd import statsmodels.formula.api as smf data = pd.DataFrame({ 'choice': ['A', 'B', 'A', 'C', 'B', 'C', 'A', 'B', 'C', 'A'], 'feature1': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10], 'feature2': [10, 9, 8, 7, 6, 5, 4, 3, 2, 1] }) # 生成整数编码与原类别映射 data['choice_code'], categories = pd.factorize(data['choice'])
2. 使用编码后的变量构建模型
把公式中的choice替换为编码后的choice_code:
model = smf.mnlogit(formula='choice_code ~ feature1 + feature2', data=data).fit()
3. 结果解读时映射回原类别
若需要对应原类别解读结果,可通过categories变量还原编码与类别的对应关系:
print(model.summary()) # 输出类别编码映射 print(f"类别编码对应: {dict(zip(range(len(categories)), categories))}")
额外说明
- 也可使用
data['choice'].astype('category').cat.codes完成编码转换,但factorize()会直接返回编码与原类别,更便于后续映射。 - 必须确保传入
mnlogit的因变量是单列整数类型,禁止使用one-hot编码后的多列数据。
内容的提问来源于stack exchange,提问作者Brian Bull
相关产品推荐
相关产品推荐

