You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLP项目语义去重:将同义重复评论替换为首次出现内容

语义相似评论去重替换方案

原代码存在的问题

  • 多层嵌套循环导致效率极低,且逻辑存在漏洞(如items被提前清空、df1['test']疑似笔误)
  • 相似度阈值(0.5-0.99)设置过于宽泛,容易造成误匹配
  • 预处理流程不统一:remove_punctuation返回句子列表,但Other_comments_unique的处理是直接转小写去标点,两者格式不一致
  • 未实现“保留首次出现内容”的核心逻辑,只是随机映射相似文本

改进方案

步骤1:统一预处理

先将评论统一处理为标准化文本(去标点、小写、词形还原),确保所有文本格式一致。

步骤2:预计算句向量

批量计算所有评论的Sentence-BERT向量,避免循环中重复编码,提升效率。

步骤3:语义相似分组

使用相似度矩阵结合阈值,或用DBSCAN聚类,将语义相似的评论归为一组。

步骤4:替换为首次出现内容

遍历每组,将组内所有评论替换为该组首次出现的原始评论内容。

改进后代码示例

import string
import pandas as pd
from nltk.tokenize import word_tokenize
from nltk.stem import WordNetLemmatizer
from sentence_transformers import SentenceTransformer, util
import numpy as np

# 统一预处理函数
def preprocess_text(text):
    if pd.isna(text) or text.strip() == "":
        return ""
    # 小写、去标点、去空格
    text = text.lower().strip().translate(str.maketrans('', '', string.punctuation))
    # 词形还原
    lemmatizer = WordNetLemmatizer()
    words = word_tokenize(text)
    lemmatized_words = [lemmatizer.lemmatize(word) for word in words]
    return " ".join(lemmatized_words)

# 加载模型
model = SentenceTransformer('paraphrase-MiniLM-L6-v2')

# 1. 预处理评论列
df1['processed_comment'] = df1["DUPLICATE - Other comments"].apply(preprocess_text)

# 2. 过滤空文本,保留有效评论并记录首次出现位置
valid_comments = df1[df1['processed_comment'] != ""].copy()
valid_comments['original_comment'] = df1["DUPLICATE - Other comments"]
valid_comments['first_occur_idx'] = valid_comments.groupby('processed_comment').index.transform('min')

# 3. 预计算所有处理后评论的句向量
embeddings = model.encode(valid_comments['processed_comment'].tolist(), convert_to_tensor=True)

# 4. 计算相似度矩阵,构建相似组
similarity_threshold = 0.75  # 可根据数据调整
cosine_scores = util.cos_sim(embeddings, embeddings)
similar_groups = {}
seen_indices = set()

for idx in valid_comments.index:
    if idx in seen_indices:
        continue
    # 筛选相似度达标索引
    similar_pos = np.where(cosine_scores[idx].cpu().numpy() >= similarity_threshold)[0]
    original_indices = valid_comments.iloc[similar_pos].index.tolist()
    # 获取组内首次出现的原始评论
    first_idx = valid_comments.loc[original_indices, 'first_occur_idx'].min()
    first_comment = valid_comments.loc[first_idx, 'original_comment']
    # 记录替换映射
    for i in original_indices:
        similar_groups[i] = first_comment
        seen_indices.add(i)

# 5. 将替换结果应用到原数据框
df1["DUPLICATE - Other comments"] = df1.apply(
    lambda row: similar_groups.get(row.name, row["DUPLICATE - Other comments"]),
    axis=1
)

额外优化建议

  • 数据量较大时,用FAISS库做近似最近邻搜索,替代全量相似度矩阵计算,大幅提升效率
  • 可更换更适配的模型,如all-MiniLM-L6-v2或多语言模型(若涉及非英文评论)
  • 通过人工标注少量样本调整相似度阈值,匹配业务场景需求

内容的提问来源于stack exchange,提问作者Bhuvaneshwari D Raman Effect

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 07:23:10