You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

RandomForestClassifier训练时特征名称类型不统一的报错解决

解决RandomForestClassifier训练时的特征名称类型不匹配问题

问题场景

处理SMSSpamCollection.tsv短信数据时,生成包含body_len、punct%特征和TF-IDF向量特征的X_features数据集,拆分训练测试集后训练RandomForestClassifier,触发如下错误:

TypeError: Feature names are only supported if all input features have string names, but your input has ['int', 'str'] as feature name / column name types.
If you want feature names to be stored and validated, you must convert them all to strings, by using X.columns = X.columns.astype(str) for example.
Otherwise you can remove feature / column names from your input data, or convert them all to a non-string data type.

错误原因

你的X_features是拼接两部分生成的:

  • 前两列来自原数据集,列名是字符串类型:'body_len'、'punct%'
  • 后面的TF-IDF特征列由稀疏矩阵转成DataFrame生成,默认列名是整数类型(0、1、2...)
    sklearn新版本对特征名称的类型一致性有严格校验,混合int和str类型的列名会触发该错误。

解决方法

方案1:统一所有列名为字符串类型(推荐)

在生成X_features后,添加一行代码将所有列名转为字符串,既满足校验要求,又能保留特征名称用于后续分析(比如查看特征重要性):

X_features = pd.concat([data['body_len'], data['punct%'], pd.DataFrame(X_tfidf.toarray())], axis=1)
# 新增该行,统一列名为字符串类型
X_features.columns = X_features.columns.astype(str)
X_features.head()

方案2:移除特征名称,直接传入数值矩阵

如果不需要保留特征名称,可以在训练时直接传入数值数组,跳过列名校验:

# 训练时用.values或.to_numpy()获取纯数值矩阵
rf_model = rf.fit(X_train.values, y_train)

原始数据处理代码

import nltk
import pandas as pd
import re
from sklearn.feature_extraction.text import TfidfVectorizer
import string

stopwords = nltk.corpus.stopwords.words('english')
ps = nltk.PorterStemmer()

data = pd.read_csv("SMSSpamCollection.tsv", sep='\t')
data.columns = ['label', 'body_text']

def count_punct(text):
    count = sum([1 for char in text if char in string.punctuation])
    return round(count/(len(text) - text.count(" ")), 3)*100

data['body_len'] = data['body_text'].apply(lambda x: len(x) - x.count(" "))
data['punct%'] = data['body_text'].apply(lambda x: count_punct(x))

def clean_text(text):
    text = "".join([word.lower() for word in text if word not in string.punctuation])
    tokens = re.split('\W+', text)
    text = [ps.stem(word) for word in tokens if word not in stopwords]
    return text

tfidf_vect = TfidfVectorizer(analyzer=clean_text)
X_tfidf = tfidf_vect.fit_transform(data['body_text'])

X_features = pd.concat([data['body_len'], data['punct%'], pd.DataFrame(X_tfidf.toarray())], axis=1)
X_features.head()

原始训练代码

from sklearn.metrics import precision_recall_fscore_support as score
from sklearn.model_selection import train_test_split


X_train, X_test, y_train, y_test = train_test_split(X_features, data['label'], test_size=0.2)


from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=50, max_depth=20, n_jobs=-1)
rf_model = rf.fit(X_train, y_train)

内容的提问来源于stack exchange,提问作者Century Egg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 17:13:11