You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将DataFrame中重复/近似重复comment替换为same并修正id?

解决DataFrame重复/近似重复comment替换及id修正问题

需求梳理

  • 将完全重复或近似重复的comment字段内容,除每组第一条外,其余替换为"same"
  • 修正id字段:每个Key分组内,按行序从1开始连续编号

实现步骤及代码

1. 依赖安装(若未安装)

首先需要安装用于文本相似度匹配的工具库:

pip install fuzzywuzzy python-Levenshtein

(注:python-Levenshtein为可选依赖,能提升相似度计算的运行效率)

2. 完整处理代码

import pandas as pd
from fuzzywuzzy import fuzz

# 构造原数据
df = {'Key': ['111', '111','111', '222*1','222*2', '333*1','333*2', '333*3','444','444', '444'],
      'id' : ['', '','', '1','2', '1','2', '3','', '','',],
      'comment': ['wrong sentence', 'wrong sentence','wrong sentence', 'M','M', 'F','F', 'F','wrong sentence used in the topic', 'wrong sentence used','wrong sentence use']}
  
df = pd.DataFrame(df)

# 步骤1:修正id字段——每个Key分组内按行号从1开始连续编号
df['id'] = df.groupby('Key').cumcount() + 1

# 步骤2:处理重复及近似重复的comment
def mark_approx_duplicates(group, threshold=80):
    # 保留每组第一条comment
    result = [group.iloc[0]['comment']]
    # 遍历组内后续条目,对比相似度
    for i in range(1, len(group)):
        current_comment = group.iloc[i]['comment']
        # 计算当前文本与组内第一条的相似度
        similarity = fuzz.ratio(group.iloc[0]['comment'], current_comment)
        # 达到相似度阈值则标记为same,否则保留原内容
        result.append('same' if similarity >= threshold else current_comment)
    group['comment'] = result
    return group

# 按Key分组处理comment字段
df = df.groupby('Key', group_keys=False).apply(mark_approx_duplicates)

print(df)

代码说明

  • id修正:通过groupby('Key').cumcount() + 1实现每个Key组内从1开始的连续编号,统一覆盖原空值或已有id
  • 近似重复判断:使用fuzz.ratio计算文本相似度(阈值设为80,可按需调整),完全重复文本相似度为100,会自动被标记为"same"
  • 分组处理:按Key分组后单独处理每组comment,确保重复/近似重复的判定仅在同组内生效

输出结果

执行代码后得到的DataFrame如下:

Key  id                          comment
0     111   1                  wrong sentence
1     111   2                            same
2     111   3                            same
3  222*1   1                                M
4  222*2   2                            same
5  333*1   1                                F
6  333*2   2                            same
7  333*3   3                            same
8     444   1  wrong sentence used in the topic
9     444   2                            same
10    444   3                            same

内容的提问来源于stack exchange,提问作者AAA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 22:55:28