查找DataFrame重复字符串并返回首个索引与重复索引对应关系
处理含重复字符串的大型DataFrame:生成首索引与重复索引的对应关系
需求:现有一个DataFrame,其中字符串内容存在重复,但对应的Index_ID不同。需要找出所有重复的字符串条目,生成首次出现的Index_ID与后续重复条目的Index_ID的对应关系,用于精简大型长字符串数据集并保留索引关联信息。
输入示例
import pandas as pd data = [[1, 'online delivery, and now offer dedicated learning platforms...'], [7, 'verything is in a state of change. There ...'], [52, 'online delivery, and now offer dedicated learning platforms...'], [84, 'verything is in a state of change. There ...'], [5699, 'online delivery, and now offer dedicated learning platforms...'], [105687, 'you have managed to get ahead'], [654684684, 'More Strings'] ] df = pd.DataFrame(data, columns=['Index_ID', 'Strings']) df.set_index('Index_ID', inplace= True)
解决方案
针对大型数据集,通过分组+筛选的方式高效实现需求,避免冗余计算:
import pandas as pd # 为每个字符串分组,标记该组首次出现的Index_ID first_index = df.groupby('Strings')['Strings'].transform(lambda x: x.index[0]) # 重置索引,方便后续筛选操作 df_reset = df.reset_index() # 筛选出所有非首次出现的重复条目 duplicates = df_reset[df_reset['Index_ID'] != first_index] # 构建目标对应关系DataFrame duplicate_indexes = pd.DataFrame({ 'First_Index_ID': first_index[duplicates.index], 'Duplicate_Index_ID': duplicates['Index_ID'] }).reset_index(drop=True)
输出结果
执行代码后得到的duplicate_indexes等价于以下构造的DataFrame:
import pandas as pd data = [[1, 52], [1, 5699], [7, 84] ] duplicate_indexes = pd.DataFrame(data, columns=['First_Index_ID', 'Duplicate_Index_ID'])
代码说明
groupby('Strings').transform:批量为每个字符串条目标记首次出现的索引,适配大型数据集的高效处理- 筛选条件
df_reset['Index_ID'] != first_index:精准定位所有重复的非首条条目 - 最终生成的DataFrame直接保留所需的索引对应关系,便于后续数据精简操作
内容的提问来源于stack exchange,提问作者Edd
相关产品推荐
相关产品推荐

