RandomForestClassifier训练时特征名称类型不统一的报错解决
问题场景
处理SMSSpamCollection.tsv短信数据时,生成包含body_len、punct%特征和TF-IDF向量特征的X_features数据集,拆分训练测试集后训练RandomForestClassifier,触发如下错误:
TypeError: Feature names are only supported if all input features have string names, but your input has ['int', 'str'] as feature name / column name types.
If you want feature names to be stored and validated, you must convert them all to strings, by using X.columns = X.columns.astype(str) for example.
Otherwise you can remove feature / column names from your input data, or convert them all to a non-string data type.
错误原因
你的X_features是拼接两部分生成的:
- 前两列来自原数据集,列名是字符串类型:
'body_len'、'punct%' - 后面的TF-IDF特征列由稀疏矩阵转成DataFrame生成,默认列名是整数类型(0、1、2...)
sklearn新版本对特征名称的类型一致性有严格校验,混合int和str类型的列名会触发该错误。
解决方法
方案1:统一所有列名为字符串类型(推荐)
在生成X_features后,添加一行代码将所有列名转为字符串,既满足校验要求,又能保留特征名称用于后续分析(比如查看特征重要性):
X_features = pd.concat([data['body_len'], data['punct%'], pd.DataFrame(X_tfidf.toarray())], axis=1) # 新增该行,统一列名为字符串类型 X_features.columns = X_features.columns.astype(str) X_features.head()
方案2:移除特征名称,直接传入数值矩阵
如果不需要保留特征名称,可以在训练时直接传入数值数组,跳过列名校验:
# 训练时用.values或.to_numpy()获取纯数值矩阵 rf_model = rf.fit(X_train.values, y_train)
原始数据处理代码
import nltk import pandas as pd import re from sklearn.feature_extraction.text import TfidfVectorizer import string stopwords = nltk.corpus.stopwords.words('english') ps = nltk.PorterStemmer() data = pd.read_csv("SMSSpamCollection.tsv", sep='\t') data.columns = ['label', 'body_text'] def count_punct(text): count = sum([1 for char in text if char in string.punctuation]) return round(count/(len(text) - text.count(" ")), 3)*100 data['body_len'] = data['body_text'].apply(lambda x: len(x) - x.count(" ")) data['punct%'] = data['body_text'].apply(lambda x: count_punct(x)) def clean_text(text): text = "".join([word.lower() for word in text if word not in string.punctuation]) tokens = re.split('\W+', text) text = [ps.stem(word) for word in tokens if word not in stopwords] return text tfidf_vect = TfidfVectorizer(analyzer=clean_text) X_tfidf = tfidf_vect.fit_transform(data['body_text']) X_features = pd.concat([data['body_len'], data['punct%'], pd.DataFrame(X_tfidf.toarray())], axis=1) X_features.head()
原始训练代码
from sklearn.metrics import precision_recall_fscore_support as score from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split(X_features, data['label'], test_size=0.2) from sklearn.ensemble import RandomForestClassifier rf = RandomForestClassifier(n_estimators=50, max_depth=20, n_jobs=-1) rf_model = rf.fit(X_train, y_train)
内容的提问来源于stack exchange,提问作者Century Egg

