如何将DataFrame中重复/近似重复comment替换为same并修正id?
解决DataFrame重复/近似重复comment替换及id修正问题
需求梳理
- 将完全重复或近似重复的
comment字段内容,除每组第一条外,其余替换为"same" - 修正
id字段:每个Key分组内,按行序从1开始连续编号
实现步骤及代码
1. 依赖安装(若未安装)
首先需要安装用于文本相似度匹配的工具库:
pip install fuzzywuzzy python-Levenshtein
(注:python-Levenshtein为可选依赖,能提升相似度计算的运行效率)
2. 完整处理代码
import pandas as pd from fuzzywuzzy import fuzz # 构造原数据 df = {'Key': ['111', '111','111', '222*1','222*2', '333*1','333*2', '333*3','444','444', '444'], 'id' : ['', '','', '1','2', '1','2', '3','', '','',], 'comment': ['wrong sentence', 'wrong sentence','wrong sentence', 'M','M', 'F','F', 'F','wrong sentence used in the topic', 'wrong sentence used','wrong sentence use']} df = pd.DataFrame(df) # 步骤1:修正id字段——每个Key分组内按行号从1开始连续编号 df['id'] = df.groupby('Key').cumcount() + 1 # 步骤2:处理重复及近似重复的comment def mark_approx_duplicates(group, threshold=80): # 保留每组第一条comment result = [group.iloc[0]['comment']] # 遍历组内后续条目,对比相似度 for i in range(1, len(group)): current_comment = group.iloc[i]['comment'] # 计算当前文本与组内第一条的相似度 similarity = fuzz.ratio(group.iloc[0]['comment'], current_comment) # 达到相似度阈值则标记为same,否则保留原内容 result.append('same' if similarity >= threshold else current_comment) group['comment'] = result return group # 按Key分组处理comment字段 df = df.groupby('Key', group_keys=False).apply(mark_approx_duplicates) print(df)
代码说明
- id修正:通过
groupby('Key').cumcount() + 1实现每个Key组内从1开始的连续编号,统一覆盖原空值或已有id - 近似重复判断:使用
fuzz.ratio计算文本相似度(阈值设为80,可按需调整),完全重复文本相似度为100,会自动被标记为"same" - 分组处理:按
Key分组后单独处理每组comment,确保重复/近似重复的判定仅在同组内生效
输出结果
执行代码后得到的DataFrame如下:
Key id comment 0 111 1 wrong sentence 1 111 2 same 2 111 3 same 3 222*1 1 M 4 222*2 2 same 5 333*1 1 F 6 333*2 2 same 7 333*3 3 same 8 444 1 wrong sentence used in the topic 9 444 2 same 10 444 3 same
内容的提问来源于stack exchange,提问作者AAA
相关产品推荐
相关产品推荐

