You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas分组DataFrame应用函数报错排查:医疗记录场景

问题分析与修正方案

你的报错"None of [Int64Index([8], dtype='int64')] are in the [index]",根源在于**x[x.location=='initial'].index+1会生成不存在于当前分组的索引**。比如ID3的分组只有一行(initial行),它的索引是8(假设原数据中该行索引为8),+1后得到9,但当前分组的索引只有8,x.loc[9]自然找不到对应的行,触发KeyError。

另外,这个写法还有潜在问题:如果分组中有多个initial行,会生成多个索引,同样可能出现索引越界的情况。

修正后的代码

def First(x):
    # 先按begin_date排序,保证时间顺序的严谨性(可选但推荐)
    x_sorted = x.sort_values('begin_date').reset_index(drop=True)
    
    # 先定位initial行,确保我们用的是出院日期
    initial_rows = x_sorted[x_sorted['location'] == 'initial']
    if initial_rows.empty:
        return 'Home'
    
    # 规则1:寻找begin_date等于initial出院日期的记录
    first_end_date = initial_rows['end_date'].iloc[0]
    match_rule1 = x_sorted.loc[x_sorted['begin_date'] == first_end_date, 'location']
    if not match_rule1.empty:
        return match_rule1.iloc[0]
    
    # 规则2:寻找initial之后的首个location
    initial_pos = initial_rows.index[0]
    after_initial = x_sorted.iloc[initial_pos + 1:]
    if not after_initial.empty:
        return after_initial['location'].iloc[0]
    
    # 规则3:前两个规则都不满足时返回Home
    return 'Home'

final = df.groupby('ID').apply(First).reset_index(name='first_site')
print(final)

关键修改点说明

  1. 排序并重置索引:
    先对分组内的数据按begin_date排序,再重置索引为连续的0、1、2...,这样后续通过位置(iloc)取数据时,不会出现索引越界的问题,同时保证时间顺序的正确性。

  2. 安全获取initial后续行:
    不再直接操作原索引,而是先找到initial行在排序后DataFrame中的位置(initial_pos),再通过iloc[initial_pos + 1:]获取后续所有行。如果initial是分组的最后一行,after_initial会是空的,自然触发规则3,不会报错。

  3. 严谨定位initial行:
    先检查分组中是否存在initial行,避免出现无initial的异常情况;同时用initial行的出院日期作为规则1的匹配依据,而不是直接取首行,更符合业务逻辑。

测试样例验证

运行修正后的代码,会得到和你预期完全一致的结果:

IDfirst_site
1rehab
2nursing
3Home
4nursing

内容的提问来源于stack exchange,提问作者CandleWax

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:07:39