NLP项目语义去重:将同义重复评论替换为首次出现内容
语义相似评论去重替换方案
原代码存在的问题
- 多层嵌套循环导致效率极低,且逻辑存在漏洞(如
items被提前清空、df1['test']疑似笔误) - 相似度阈值(0.5-0.99)设置过于宽泛,容易造成误匹配
- 预处理流程不统一:
remove_punctuation返回句子列表,但Other_comments_unique的处理是直接转小写去标点,两者格式不一致 - 未实现“保留首次出现内容”的核心逻辑,只是随机映射相似文本
改进方案
步骤1:统一预处理
先将评论统一处理为标准化文本(去标点、小写、词形还原),确保所有文本格式一致。
步骤2:预计算句向量
批量计算所有评论的Sentence-BERT向量,避免循环中重复编码,提升效率。
步骤3:语义相似分组
使用相似度矩阵结合阈值,或用DBSCAN聚类,将语义相似的评论归为一组。
步骤4:替换为首次出现内容
遍历每组,将组内所有评论替换为该组首次出现的原始评论内容。
改进后代码示例
import string import pandas as pd from nltk.tokenize import word_tokenize from nltk.stem import WordNetLemmatizer from sentence_transformers import SentenceTransformer, util import numpy as np # 统一预处理函数 def preprocess_text(text): if pd.isna(text) or text.strip() == "": return "" # 小写、去标点、去空格 text = text.lower().strip().translate(str.maketrans('', '', string.punctuation)) # 词形还原 lemmatizer = WordNetLemmatizer() words = word_tokenize(text) lemmatized_words = [lemmatizer.lemmatize(word) for word in words] return " ".join(lemmatized_words) # 加载模型 model = SentenceTransformer('paraphrase-MiniLM-L6-v2') # 1. 预处理评论列 df1['processed_comment'] = df1["DUPLICATE - Other comments"].apply(preprocess_text) # 2. 过滤空文本,保留有效评论并记录首次出现位置 valid_comments = df1[df1['processed_comment'] != ""].copy() valid_comments['original_comment'] = df1["DUPLICATE - Other comments"] valid_comments['first_occur_idx'] = valid_comments.groupby('processed_comment').index.transform('min') # 3. 预计算所有处理后评论的句向量 embeddings = model.encode(valid_comments['processed_comment'].tolist(), convert_to_tensor=True) # 4. 计算相似度矩阵,构建相似组 similarity_threshold = 0.75 # 可根据数据调整 cosine_scores = util.cos_sim(embeddings, embeddings) similar_groups = {} seen_indices = set() for idx in valid_comments.index: if idx in seen_indices: continue # 筛选相似度达标索引 similar_pos = np.where(cosine_scores[idx].cpu().numpy() >= similarity_threshold)[0] original_indices = valid_comments.iloc[similar_pos].index.tolist() # 获取组内首次出现的原始评论 first_idx = valid_comments.loc[original_indices, 'first_occur_idx'].min() first_comment = valid_comments.loc[first_idx, 'original_comment'] # 记录替换映射 for i in original_indices: similar_groups[i] = first_comment seen_indices.add(i) # 5. 将替换结果应用到原数据框 df1["DUPLICATE - Other comments"] = df1.apply( lambda row: similar_groups.get(row.name, row["DUPLICATE - Other comments"]), axis=1 )
额外优化建议
- 数据量较大时,用FAISS库做近似最近邻搜索,替代全量相似度矩阵计算,大幅提升效率
- 可更换更适配的模型,如
all-MiniLM-L6-v2或多语言模型(若涉及非英文评论) - 通过人工标注少量样本调整相似度阈值,匹配业务场景需求
内容的提问来源于stack exchange,提问作者Bhuvaneshwari D Raman Effect
相关产品推荐
相关产品推荐

