You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LogisticRegression文本分类器合并特征名触发ValueError问题求助

问题解决:合并文本特征与自定义特征的特征名

报错原因

vectorizer.get_feature_names_out()返回的是numpy字符串数组,直接用+和Python列表拼接时,numpy会将该操作解析为数组元素级运算,而两个对象的形状(15569,)和(3,)不匹配,因此触发广播错误。另外,自定义特征是你额外添加的,CountVectorizer不会自动将其纳入自身的特征名列表,所以必须手动合并两部分名称。

正确处理方式

有两种可靠的合并方法:

方法1:转为Python列表后拼接

先将向量器生成的特征名转为列表,再与自定义特征名列表拼接,最后按需转为numpy数组:

# 提取文本特征名并转为列表
text_feature_names = vectorizer.get_feature_names_out().tolist()
# 定义自定义特征名列表
custom_feature_names = ['length_of_doc', 'digit_count', 'non_word_char_count']
# 合并后转为numpy数组(如果需要)
feature_names = np.array(text_feature_names + custom_feature_names)

方法2:使用numpy的concatenate函数拼接数组

将自定义特征名转为numpy数组,再用concatenate完成拼接:

text_features = vectorizer.get_feature_names_out()
custom_features = np.array(['length_of_doc', 'digit_count', 'non_word_char_count'])
# 拼接两个numpy数组
feature_names = np.concatenate([text_features, custom_features])

内容的提问来源于stack exchange,提问作者Ubaidul Khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 15:41:46