为什么pandas的for循环中不能用loc(i+1,'列名'),出现KeyError如何解决?
报错原因
- 循环终止条件设置错误:你的DataFrame总行数为
pdf.shape[0],若使用默认从0开始的连续整数索引,最大合法索引值为pdf.shape[0]-1。当循环执行到最后一行时,i的取值为pdf.shape[0]-1,此时i < pdf.shape[0]判定为真,执行pdf.loc[i+1,'A']时会访问不存在的索引,触发KeyError。 - 索引兼容问题:如果你的DataFrame做过行删除、过滤等操作,索引并非连续整数,直接用
i+1通过loc取值也会找不到对应行,触发报错。
修复方案
1. 修复原相邻行比较逻辑
推荐使用位置索引iloc取值,避免索引不连续的问题,同时调整循环范围:
# 仅遍历到倒数第二行,避免访问超出范围的行 for idx in range(pdf.shape[0] - 1): current_a = pdf.iloc[idx][pdf.columns.get_loc('A')] next_a = pdf.iloc[idx + 1][pdf.columns.get_loc('A')] if next_a != current_a: print("a") else: print("b") print("finished")
2. 实现相邻行相乘加3的计算逻辑
按你需要的计算规则,循环从第二行开始遍历即可:
# idx从1开始,对应第二行,依次和前一行计算 for idx in range(1, pdf.shape[0]): prev_a = pdf.iloc[idx - 1][pdf.columns.get_loc('A')] current_a = pdf.iloc[idx][pdf.columns.get_loc('A')] print(current_a * prev_a + 3)
如果不需要强制用for循环,用pandas内置的shift方法性能更高,代码更简洁:
# shift(1)将A列所有值向下挪一行,自动和前一行对齐 calc_result = pdf['A'] * pdf['A'].shift(1) + 3 # 输出所有非空的计算结果(第一行没有前一行,结果为空) print(calc_result.dropna().to_list())
内容的提问来源于stack exchange,提问作者Ryou
相关产品推荐
相关产品推荐

