Pandas使用set_index后存在重复索引时如何获取首个匹配值
解决方案
方式1:直接取第一个值(兼容重复/非重复场景)
不需要额外检测,直接通过.iloc[0]固定取首个匹配值,改动最小:
teams = df_user.set_index(['name','Email'])["Team"] # 用元组形式传多重索引到loc避免链式索引问题,iloc[0]取第一个值 author_team = teams.loc[(name, email)].iloc[0] return author_team
方式2:先检测是否存在多个值,再按需求取值
如果需要针对多值场景做特殊处理,可以先判断返回结果的长度:
teams = df_user.set_index(['name','Email'])["Team"] match_res = teams.loc[(name, email)] if len(match_res) > 1: # 存在多个匹配值,可按需取第n个,取第二个就换成iloc[1] author_team = match_res.iloc[0] else: # 仅单个匹配值,直接返回标量 author_team = match_res.item() return author_team
方式3:提前对索引去重,后续查询无需额外处理
如果所有查询都默认取第一个匹配值,可以在构造teams的时候就对重复索引做去重,后续查询直接取即可:
# 按索引分组,每组保留第一个值 teams = df_user.set_index(['name','Email'])["Team"].groupby(level=[0,1]).first() author_team = teams.loc[(name, email)] return author_team
内容的提问来源于stack exchange,提问作者Andrew D
相关产品推荐
相关产品推荐

