You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用FastText输出得到稀疏矩阵以训练机器学习模型

FastText生成训练用向量的实现方法

你当前调用的most_similar接口作用是返回与输入文本/词语语义最接近的词条及相似度,并非用于生成训练所需的向量格式,要实现你的需求可以按以下步骤操作:

第一步:获取FastText稠密向量

FastText原生支持直接输出词语和句子的稠密向量,可直接用于机器学习模型训练:

  • 单个词语向量获取:直接通过词表索引取值
# 获取词语school的向量,返回维度为你训练时设置的vector_size(如100、200)
word_vec = model_ted.wv['school']
  • 长文本句向量获取:调用对应接口直接生成整段文本的语义向量
long_text = "The Lemon Drop Kid , a New York City swindler, is illegally touting horses at a Florida racetrack. After several successful hustles, the Kid comes across a beautiful, but gullible, woman intending to bet a lot of money. The Kid convinces her to switch her bet, employing a prefabricated con. Unfortunately for the Kid, the woman belongs to notorious gangster Moose Moran , as does the money. The Kid's choice finishes dead last and a furious Moran demands the Kid provide him with $10,000  by Christmas Eve, or the Kid won't make it to New Year's. The Kid decides to return to New York to try to come up with the money. He first tries his on-again, off-again girlfriend Brainy Baxter . However, when talk of long-term commitment arises, the Kid quickly makes an escape."
# 生成整段文本的句向量
sentence_vec = model_ted.get_sentence_vector(long_text)

注:如果你使用的是gensim库训练的FastText模型,句向量可通过model_ted.wv.get_mean_vector(long_text.split())生成。

第二步:转换为CSR稀疏矩阵(和TF-IDF输出格式对齐)

如果你需要和之前TF-IDF输出的1xN稀疏矩阵格式完全一致,可以通过scipy工具转换:

import numpy as np
from scipy.sparse import csr_matrix

# 示例生成1x10000维度的稀疏矩阵
target_dim = 10000
# 若当前向量维度小于目标维度,先补全维度
if len(sentence_vec) < target_dim:
    sentence_vec = np.pad(sentence_vec, (0, target_dim - len(sentence_vec)), mode='constant')
# 可设置阈值过滤小数值,提升稀疏度
threshold = 1e-3
sentence_vec[np.abs(sentence_vec) < threshold] = 0
# 转换为CSR格式稀疏矩阵
sparse_vec = csr_matrix(sentence_vec.reshape(1, -1))
# 输出格式即符合你要求的<1x10000 sparse matrix of type '<class 'numpy.float64'>' with x stored elements in Compressed Sparse Row format>
print(sparse_vec)

第三步:用于机器学习模型训练

生成的稠密向量或转换后的稀疏矩阵,可直接替代之前的TF-IDF矩阵输入到机器学习模型中,无需修改原有训练逻辑:

  • 批量处理所有训练文本,每条文本生成对应向量,拼接为 shape 为(样本数, 向量维度)的矩阵
  • 直接将矩阵与对应标签传入模型的fit方法即可完成训练

内容的提问来源于stack exchange,提问作者Joseph

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 20:54:01