You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

查找DataFrame重复字符串并返回首个索引与重复索引对应关系

处理含重复字符串的大型DataFrame:生成首索引与重复索引的对应关系

需求:现有一个DataFrame,其中字符串内容存在重复,但对应的Index_ID不同。需要找出所有重复的字符串条目,生成首次出现的Index_ID与后续重复条目的Index_ID的对应关系,用于精简大型长字符串数据集并保留索引关联信息。

输入示例

import pandas as pd
data = [[1, 'online delivery, and now offer dedicated learning platforms...'], 
        [7, 'verything is in a state of change. There ...'], 
        [52, 'online delivery, and now offer dedicated learning platforms...'],
        [84, 'verything is in a state of change. There ...'],
        [5699, 'online delivery, and now offer dedicated learning platforms...'],
        [105687, 'you have managed to get ahead'],
        [654684684, 'More Strings']
        ]
  
df = pd.DataFrame(data, columns=['Index_ID', 'Strings'])
df.set_index('Index_ID', inplace= True)

解决方案

针对大型数据集,通过分组+筛选的方式高效实现需求,避免冗余计算:

import pandas as pd

# 为每个字符串分组,标记该组首次出现的Index_ID
first_index = df.groupby('Strings')['Strings'].transform(lambda x: x.index[0])

# 重置索引,方便后续筛选操作
df_reset = df.reset_index()

# 筛选出所有非首次出现的重复条目
duplicates = df_reset[df_reset['Index_ID'] != first_index]

# 构建目标对应关系DataFrame
duplicate_indexes = pd.DataFrame({
    'First_Index_ID': first_index[duplicates.index],
    'Duplicate_Index_ID': duplicates['Index_ID']
}).reset_index(drop=True)

输出结果

执行代码后得到的duplicate_indexes等价于以下构造的DataFrame:

import pandas as pd
data = [[1, 52], 
        [1, 5699],
        [7, 84]        
       ]
  
duplicate_indexes = pd.DataFrame(data, columns=['First_Index_ID', 'Duplicate_Index_ID'])

代码说明

  • groupby('Strings').transform:批量为每个字符串条目标记首次出现的索引,适配大型数据集的高效处理
  • 筛选条件df_reset['Index_ID'] != first_index:精准定位所有重复的非首条条目
  • 最终生成的DataFrame直接保留所需的索引对应关系,便于后续数据精简操作

内容的提问来源于stack exchange,提问作者Edd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 11:55:25