使用groupby后列乘标量报错:索引不兼容问题求助
解决TypeError: incompatible index of inserted column with frame index错误
问题场景
代码:
tmp_ml['mode_dur_secs_lb'] = tmp_ml.groupby(['id', 'c_num']).apply( lambda x: x['mode_duration_secs']*0.9)
数据表格:
| id | c_num | mode_duration_secs |
|---|---|---|
| a1 | 116 | 20 |
| a1 | 279 | 3 |
| a2 | 9 | 19 |
| a3 | 16 | 16 |
| a3 | 17 | 19 |
运行后抛出错误:TypeError: incompatible index of inserted column with frame index
错误原因
用groupby(['id', 'c_num']).apply()处理后,返回的Series会带上多层索引(MultiIndex),包含分组的id、c_num以及原数据的行索引。但原DataFrametmp_ml用的是默认单级整数索引,两者索引结构不匹配,导致赋值时出错。
解决方案
两种简单修复方式:
方式1:用transform替代apply
transform会自动保留原DataFrame的索引结构,直接返回和原数据行数匹配的结果,适合这种分组内逐行做简单运算的场景:
tmp_ml['mode_dur_secs_lb'] = tmp_ml.groupby(['id', 'c_num'])['mode_duration_secs'].transform(lambda x: x*0.9)
方式2:重置apply结果的索引
如果一定要用apply,可以调用reset_index(drop=True)去掉多余的分组索引,让结果索引和原DataFrame对齐:
tmp_ml['mode_dur_secs_lb'] = tmp_ml.groupby(['id', 'c_num']).apply( lambda x: x['mode_duration_secs']*0.9).reset_index(drop=True)
验证结果
修复后,tmp_ml会新增正确计算的列:
| id | c_num | mode_duration_secs | mode_dur_secs_lb |
|---|---|---|---|
| a1 | 116 | 20 | 18.0 |
| a1 | 279 | 3 | 2.7 |
| a2 | 9 | 19 | 17.1 |
| a3 | 16 | 16 | 14.4 |
| a3 | 17 | 19 | 17.1 |
内容的提问来源于stack exchange,提问作者user7446890
相关产品推荐
相关产品推荐

