You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用sklearn训练菜系预测模型时CSV字符串列转浮点报错如何解决

报错原因

你遇到的ValueError: could not convert string to float: 'indian'报错,是因为sklearn中的机器学习模型仅支持数值类型输入,而你输入的B列是字符串类型的分类值(从报错内容看属于菜系分类值),无法被自动转换为浮点型数值输入模型。

解决方案

首先确认B列的作用,对应选择处理方式:

场景1:B列为待预测的目标标签(y)

使用标签编码转换分类值为数值,预测后可反向转回原菜系名称:

from sklearn.preprocessing import LabelEncoder
import pandas as pd

# 读取csv数据
df = pd.read_csv("你的训练数据路径.csv")
label_encoder = LabelEncoder()
# 对B列(假设列名为cuisine,可替换为你实际的列名)做编码
y = label_encoder.fit_transform(df["cuisine"])

# 模型预测后,可通过如下方法将数值预测结果转回原菜系名
pred_cuisine_names = label_encoder.inverse_transform(trained_model.predict(X_test))

场景2:B列为输入特征的一部分

根据特征属性选择编码方式:

  • 若为无序分类特征(菜系分类无高低等级之分,该场景下多用该方案),使用独热编码:
from sklearn.preprocessing import OneHotEncoder

onehot_encoder = OneHotEncoder(sparse_output=False, drop="first")
# 对B列做独热编码
b_encoded = onehot_encoder.fit_transform(df[["cuisine"]])
# 将编码后的特征和其他输入特征合并
X = pd.concat(
    [df.drop("cuisine", axis=1).reset_index(drop=True), 
     pd.DataFrame(b_encoded, columns=onehot_encoder.get_feature_names_out(["cuisine"]))],
    axis=1
)
  • 若为有序分类特征(存在明确的等级排序,菜系预测场景基本不适用),可使用OrdinalEncoder做编码。
注意事项
  • 编码器仅用训练集数据拟合,再对测试集做变换,避免出现数据泄露问题
  • 不要直接对全量数据集做编码后再拆分训练测试集,会导致模型泛化能力下降

内容的提问来源于stack exchange,提问作者Joel K Nyongesa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 15:09:03