如何在Python的TF-IDF Vectorizer中分组词汇以缩减词汇量?
解决服装数据集TF-IDF词汇分组与降维问题
针对你44000条服装句子的TF-IDF矩阵规模过大、计算耗时的问题,完全可以通过词汇分组映射来合并同类词汇,让它们共享相同的TF-IDF值,同时缩减矩阵维度。以下是几种实用方案,结合你的现有代码修改:
方法一:预处理阶段做词汇替换(最直观)
先定义一个词汇映射字典,把同类词统一替换成目标词,再对处理后的文本生成TF-IDF矩阵。这种方式逻辑清晰,适合你明确知道分组规则的场景。
import pandas as pd from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.metrics.pairwise import cosine_similarity # 1. 定义词汇分组映射字典 vocab_mapping = { # 颜色分组 'teal': 'blue', 'navy': 'blue', 'turquoise': 'blue', 'aqua': 'blue', # 上衣分组 't-shirt': 'shirt', 'sweatshirt': 'shirt', 'polo': 'shirt', 'blouse': 'shirt' # 可根据需求继续添加其他分组 } # 2. 编写替换函数 def replace_vocab(text): words = text.lower().split() replaced_words = [vocab_mapping.get(word, word) for word in words] return ' '.join(replaced_words) # 3. 读取并处理数据 dataset_2 = "/dataset_files/styles_2.csv" df = pd.read_csv(dataset_2) df = df.drop(['gender', 'masterCategory', 'subCategory', 'articleType', 'baseColour', 'season', 'year', 'usage'], axis=1) # 4. 对文本做词汇替换 df['ProcessedDisplayName'] = df['ProductDisplayName'].apply(replace_vocab) # 5. 生成TF-IDF矩阵(此时同类词已被统一) tfidf = TfidfVectorizer(stop_words='english') tfidf_matrix = tfidf.fit_transform(df['ProcessedDisplayName']) cos_sim = cosine_similarity(tfidf_matrix, tfidf_matrix)
方法二:自定义TF-IDF的Tokenizer(更灵活)
直接在TF-IDF的分词阶段完成词汇映射,不需要单独预处理文本。这种方式适合把映射逻辑整合到向量生成流程中:
import pandas as pd from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.metrics.pairwise import cosine_similarity vocab_mapping = { 'teal': 'blue', 'navy': 'blue', 'turquoise': 'blue', 't-shirt': 'shirt', 'sweatshirt': 'shirt' } # 自定义tokenizer函数 def custom_tokenizer(text): words = text.lower().split() return [vocab_mapping.get(word, word) for word in words] # 读取数据 dataset_2 = "/dataset_files/styles_2.csv" df = pd.read_csv(dataset_2) df = df.drop(['gender', 'masterCategory', 'subCategory', 'articleType', 'baseColour', 'season', 'year', 'usage'], axis=1) # 使用自定义tokenizer初始化TF-IDF tfidf = TfidfVectorizer(stop_words='english', tokenizer=custom_tokenizer) tfidf_matrix = tfidf.fit_transform(df['ProductDisplayName']) cos_sim = cosine_similarity(tfidf_matrix, tfidf_matrix)
额外优化建议
- 如果词汇量还是过大,可以通过
TfidfVectorizer的max_features参数限制保留的高频词汇数量,比如max_features=5000,只保留TF-IDF权重最高的5000个词。 - 对于服装领域的通用同义词,也可以结合NLTK的同义词库(如WordNet)做批量映射,但需要手动过滤无关同义词,避免误替换。
内容的提问来源于stack exchange,提问作者rainxx
相关产品推荐
相关产品推荐

