You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scikit-learn中结合TF-IDF矩阵与Numpy数组构建联合特征矩阵?

问题原因

你对FeatureUnion(以及make_union)的核心用法理解错了:它的作用是组合特征转换的流程,接收的是转换器对象(比如TfidfVectorizer这种能完成fit/transform操作的类实例),而不是已经转换好的特征矩阵。你直接把生成后的representation1、representation2传进去,这些是数组/稀疏矩阵,不是转换器,自然会抛出没有get_feature_names_out方法的错误。

另外顺便提一句,你代码里count_vectorizer.fit_transform(input)执行了两次,完全没必要,应该先fit一次,再transform两次,节省计算资源:

count_vectorizer = CountVectorizer()
count_matrix = count_vectorizer.fit_transform(input).toarray()
sum_words = np.sum(count_matrix, axis=-1)
sum_different_words = np.count_nonzero(count_matrix, axis=-1)
representation2 = np.divide(sum_different_words, sum_words)
两种解决方案

方案一:直接合并已生成的特征矩阵(最简单)

既然你已经得到了两个特征矩阵,直接用scipy.sparse.hstack合并就行(因为representation1是稀疏矩阵,representation2是一维数组,需要先转成二维再合并):

from scipy.sparse import hstack

# 将representation2转成二维数组,匹配representation1的维度
representation2_2d = representation2.reshape(-1, 1)
# 合并稀疏矩阵和稠密数组
representation = hstack([representation1, representation2_2d])

如果需要获取特征名,可以手动拼接:

tfidf_feature_names = vectorizer.get_feature_names_out()
custom_feature_names = ["unique_word_ratio"]
all_feature_names = np.concatenate([tfidf_feature_names, custom_feature_names])

方案二:用make_union+FunctionTransformer包装自定义特征(符合sklearn流水线规范)

如果你想遵循sklearn的流水线流程,把整个特征生成逻辑统一管理,需要把自定义的特征计算逻辑包装成FunctionTransformer,然后和TfidfVectorizer一起传给make_union:

from sklearn.preprocessing import FunctionTransformer
from sklearn.pipeline import make_union

# 定义计算自定义特征的函数
def calculate_unique_word_ratio(texts):
    count_vectorizer = CountVectorizer()
    count_matrix = count_vectorizer.fit_transform(texts).toarray()
    sum_words = np.sum(count_matrix, axis=-1)
    sum_different_words = np.count_nonzero(count_matrix, axis=-1)
    # 返回二维数组,符合转换器输出要求
    return np.divide(sum_different_words, sum_words).reshape(-1, 1)

# 创建转换器:TF-IDF转换器 + 自定义特征转换器
tfidf_transformer = TfidfVectorizer(lowercase=False)
custom_transformer = FunctionTransformer(calculate_unique_word_ratio)

# 组合转换器
union = make_union(tfidf_transformer, custom_transformer)

# 生成合并后的特征矩阵
representation = union.fit_transform(input)

# 获取所有特征名
all_feature_names = union.get_feature_names_out()

注意:这里calculate_unique_word_ratio里的CountVectorizer每次调用都会重新fit,如果想和之前的TF-IDF用同一个词汇表,可以把CountVectorizer的vocab设置成TfidfVectorizer的vocab,避免不一致。

内容的提问来源于stack exchange,提问作者TiMauzi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 00:05:17