You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python的TF-IDF Vectorizer中分组词汇以缩减词汇量?

解决服装数据集TF-IDF词汇分组与降维问题

针对你44000条服装句子的TF-IDF矩阵规模过大、计算耗时的问题,完全可以通过词汇分组映射来合并同类词汇,让它们共享相同的TF-IDF值,同时缩减矩阵维度。以下是几种实用方案,结合你的现有代码修改:

方法一:预处理阶段做词汇替换(最直观)

先定义一个词汇映射字典,把同类词统一替换成目标词,再对处理后的文本生成TF-IDF矩阵。这种方式逻辑清晰,适合你明确知道分组规则的场景。

import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

# 1. 定义词汇分组映射字典
vocab_mapping = {
    # 颜色分组
    'teal': 'blue',
    'navy': 'blue',
    'turquoise': 'blue',
    'aqua': 'blue',
    # 上衣分组
    't-shirt': 'shirt',
    'sweatshirt': 'shirt',
    'polo': 'shirt',
    'blouse': 'shirt'
    # 可根据需求继续添加其他分组
}

# 2. 编写替换函数
def replace_vocab(text):
    words = text.lower().split()
    replaced_words = [vocab_mapping.get(word, word) for word in words]
    return ' '.join(replaced_words)

# 3. 读取并处理数据
dataset_2 = "/dataset_files/styles_2.csv"
df = pd.read_csv(dataset_2)
df = df.drop(['gender', 'masterCategory', 'subCategory', 'articleType', 'baseColour', 'season', 'year', 'usage'], axis=1)

# 4. 对文本做词汇替换
df['ProcessedDisplayName'] = df['ProductDisplayName'].apply(replace_vocab)

# 5. 生成TF-IDF矩阵(此时同类词已被统一)
tfidf = TfidfVectorizer(stop_words='english') 
tfidf_matrix = tfidf.fit_transform(df['ProcessedDisplayName'])
cos_sim = cosine_similarity(tfidf_matrix, tfidf_matrix)

方法二:自定义TF-IDF的Tokenizer(更灵活)

直接在TF-IDF的分词阶段完成词汇映射,不需要单独预处理文本。这种方式适合把映射逻辑整合到向量生成流程中:

import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

vocab_mapping = {
    'teal': 'blue', 'navy': 'blue', 'turquoise': 'blue',
    't-shirt': 'shirt', 'sweatshirt': 'shirt'
}

# 自定义tokenizer函数
def custom_tokenizer(text):
    words = text.lower().split()
    return [vocab_mapping.get(word, word) for word in words]

# 读取数据
dataset_2 = "/dataset_files/styles_2.csv"
df = pd.read_csv(dataset_2)
df = df.drop(['gender', 'masterCategory', 'subCategory', 'articleType', 'baseColour', 'season', 'year', 'usage'], axis=1)

# 使用自定义tokenizer初始化TF-IDF
tfidf = TfidfVectorizer(stop_words='english', tokenizer=custom_tokenizer) 
tfidf_matrix = tfidf.fit_transform(df['ProductDisplayName'])
cos_sim = cosine_similarity(tfidf_matrix, tfidf_matrix)

额外优化建议

  • 如果词汇量还是过大,可以通过TfidfVectorizer的max_features参数限制保留的高频词汇数量,比如max_features=5000,只保留TF-IDF权重最高的5000个词。
  • 对于服装领域的通用同义词,也可以结合NLTK的同义词库(如WordNet)做批量映射,但需要手动过滤无关同义词,避免误替换。

内容的提问来源于stack exchange,提问作者rainxx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 21:15:36