You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将列转为category类型后,随机森林训练报错:无法将字符串转为float

问题解答:将列转为category后随机森林仍报字符串转float错误

原因

Pandas的category类型只是给字符串列添加了类别标签,内部虽存储了编码,但scikit-learn的RandomForestClassifier不会自动解析这种类型的编码逻辑,仍会将该列视为普通字符串列。而随机森林模型仅支持数值型输入,因此会抛出"could not convert string to float"的错误。

解决方法

1. 提取category的整数编码

直接将category类型转为对应的唯一整数编码,让模型可以识别:

df["Geography"] = df["Geography"].astype("category").cat.codes

2. 独热编码(适合无顺序的类别)

如果类别之间没有先后逻辑,推荐用独热编码生成二进制特征列,避免引入无关的顺序关系:

import pandas as pd
df = pd.get_dummies(df, columns=["Geography"], drop_first=True)

3. 拆分数据集后统一编码(避免数据泄露)

如果已拆分训练集和测试集,必须基于训练集的类别信息转换测试集,保证编码规则一致:

from sklearn.preprocessing import LabelEncoder
from sklearn.model_selection import train_test_split

X = df.drop('target', axis=1)
y = df['target']
X_train, X_test, y_train, y_test = train_test_split(X,y, test_size=0.15, random_state=123, stratify=y)

le = LabelEncoder()
# 仅在训练集上拟合编码器
X_train["Geography"] = le.fit_transform(X_train["Geography"])
# 用训练集的编码器转换测试集
X_test["Geography"] = le.transform(X_test["Geography"])

内容的提问来源于stack exchange,提问作者gosiia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 02:00:19