You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用sklearn的Logistic Regression训练泰坦尼克数据集遇TypeError问题

解决LogisticRegression.fit()触发的TypeError问题

问题原因

你的特征列名同时包含整数类型和字符串类型,新版本scikit-learn对输入特征的名称类型一致性有严格校验,因此调用fit()时触发了类型错误。

解决方案

提供三种可行的修复方式,选择最适合你的场景即可:

方法1:将所有列名转为字符串类型(推荐)

在划分数据集前,把X的所有列名统一转成字符串,既解决类型问题,又保留列名信息:

X = titanic_data.drop("Survived", axis=1)
# 统一转换列名为字符串
X.columns = X.columns.astype(str)
y = titanic_data['Survived']

# 后续训练代码保持不变
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.33, random_state=42)

from sklearn.linear_model import LogisticRegression
logmodel = LogisticRegression()
logmodel.fit(X_train, y_train)

方法2:移除列名(转为numpy数组)

如果不需要保留列名信息,可以直接把DataFrame转为numpy数组,避免传递列名给模型:

# 转为numpy数组,丢弃列名
X = titanic_data.drop("Survived", axis=1).values
y = titanic_data['Survived'].values

from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.33, random_state=42)

from sklearn.linear_model import LogisticRegression
logmodel = LogisticRegression()
logmodel.fit(X_train, y_train)

方法3:将所有列名转为整数类型(仅适用特定场景)

如果你的列名原本可安全转换为整数(比如部分列名是数字字符串),可以统一转为整数类型:

X = titanic_data.drop("Survived", axis=1)
# 统一转换列名为整数(注意:若列名包含非数字字符串会报错)
X.columns = X.columns.astype(int)
y = titanic_data['Survived']

# 后续训练代码保持不变

内容的提问来源于stack exchange,提问作者Kartick Upadhyay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 22:11:06