You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用RandomForestClassifier预测未知分类变量时遇类型转换错误

解决RandomForestClassifier训练时"could not convert string to float"错误

问题背景

尝试使用RandomForestClassifier预测数据集里未知的smoking_status分类变量,但运行代码时触发如下错误:

ValueError: could not convert string to float: 'Male'

错误原因

  1. 分类特征未编码:Scikit-learn的机器学习模型仅支持数值型输入,训练数据X_train中包含gender、ever_married、work_type、Residence_type等字符串类型的分类特征,未做数值转换。
  2. 代码逻辑问题:
    • 提前使用了未定义的x_columns_to_drop、y_columns_to_drop变量;
    • 多余的train_test_split代码被后续赋值覆盖,无实际作用;
    • 最后赋值时列名拼写错误(smokingstatus应为smoking_status)。

修复步骤

  • 对所有字符串分类特征进行编码,使用LabelEncoder(适合二分类或有序分类)或OneHotEncoder(适合无序多分类),注意仅在训练集上拟合编码器,避免数据泄露;
  • 调整变量定义顺序,确保使用前先定义;
  • 移除无效的train_test_split代码;
  • 修正列名拼写错误。

完整修正代码

import pandas as pd
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.preprocessing import LabelEncoder

# 加载数据
raw_data = pd.read_csv('/Users/name/Desktop/CompSci_Projects/Stroke_Prediction_Model/dataset/healthcare-dataset-stroke-data.csv')
data = raw_data.copy(deep=True)
data.interpolate(method='linear', inplace=True)

# 先定义需要丢弃的列
x_columns_to_drop = ['id', 'smoking_status']
y_columns_to_drop = ['id', 'gender', 'age', 'hypertension', 'heart_disease', 
                     'ever_married', 'work_type', 'Residence_type', 'avg_glucose_level', 'bmi', 'stroke']

# 拆分已知/未知smoking_status的数据集
smokestatus_known = data[data['smoking_status'] != 'Unknown']
smokestatus_unknown = data[data['smoking_status'] == 'Unknown']

# 准备训练和测试数据
X_train = smokestatus_known.drop(x_columns_to_drop, axis=1)
Y_train = smokestatus_known.drop(y_columns_to_drop, axis=1)
X_test = smokestatus_unknown.drop(x_columns_to_drop, axis=1)

# 对分类特征进行LabelEncoder编码
categorical_cols = ['gender', 'ever_married', 'work_type', 'Residence_type']
encoders = {}
for col in categorical_cols:
    le = LabelEncoder()
    # 仅在训练集上拟合
    X_train[col] = le.fit_transform(X_train[col])
    # 用训练集的编码器转换测试集
    X_test[col] = le.transform(X_test[col])
    encoders[col] = le

# 训练模型
classifier = RandomForestClassifier(random_state=42)
classifier.fit(X_train, Y_train.values.ravel())  # ravel()避免DataFrame格式问题

# 预测并填充未知值
predicted_values = classifier.predict(X_test)
data.loc[data['smoking_status'] == 'Unknown', 'smoking_status'] = predicted_values

print(data)

补充说明

  • 如果是无序多分类特征(比如work_type),更推荐使用OneHotEncoder或pd.get_dummies,避免模型误认为特征存在顺序关系;
  • 编码时必须保证训练集和测试集使用相同的编码器规则,否则会导致特征分布不一致,影响预测结果。

内容的提问来源于stack exchange,提问作者rts2027

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 10:52:06