LogisticRegression文本分类器合并特征名触发ValueError问题求助
问题解决:合并文本特征与自定义特征的特征名
报错原因
vectorizer.get_feature_names_out()返回的是numpy字符串数组,直接用+和Python列表拼接时,numpy会将该操作解析为数组元素级运算,而两个对象的形状(15569,)和(3,)不匹配,因此触发广播错误。另外,自定义特征是你额外添加的,CountVectorizer不会自动将其纳入自身的特征名列表,所以必须手动合并两部分名称。
正确处理方式
有两种可靠的合并方法:
方法1:转为Python列表后拼接
先将向量器生成的特征名转为列表,再与自定义特征名列表拼接,最后按需转为numpy数组:
# 提取文本特征名并转为列表 text_feature_names = vectorizer.get_feature_names_out().tolist() # 定义自定义特征名列表 custom_feature_names = ['length_of_doc', 'digit_count', 'non_word_char_count'] # 合并后转为numpy数组(如果需要) feature_names = np.array(text_feature_names + custom_feature_names)
方法2:使用numpy的concatenate函数拼接数组
将自定义特征名转为numpy数组,再用concatenate完成拼接:
text_features = vectorizer.get_feature_names_out() custom_features = np.array(['length_of_doc', 'digit_count', 'non_word_char_count']) # 拼接两个numpy数组 feature_names = np.concatenate([text_features, custom_features])
内容的提问来源于stack exchange,提问作者Ubaidul Khan
相关产品推荐
相关产品推荐

