访问pandas DataFrame元素引发KeyError崩溃问题咨询
问题根源
- 你对
df_alter['hypothesis']这个Series用[idx]访问时,pandas默认优先匹配行索引标签,而不是行的位置顺序。你之前大概率对这个DataFrame做过行删除、过滤、取子集等操作,导致行索引不再是连续的0、1、2……n-1的整数序列,比如索引可能变成了[0,2,3,4……],当循环到idx=1时,找不到标签为1的行,就抛出了KeyError。 - 额外需要注意的是,你用的
df_alter['hypothesis'][idx] = xxx属于链式赋值,是pandas不推荐的写法,极易触发SettingWithCopyWarning,甚至可能出现修改后的值没有真正写入原DataFrame的问题。
解决方法
方法1:最小改动修复现有循环
用iloc明确指定按位置访问即可,不需要改动其他逻辑:
import re print(len(df_alter['hypothesis'])) for idx in range(len(df_alter['hypothesis'])): print(df_alter['hypothesis'].iloc[idx]) df_alter['hypothesis'].iloc[idx] = re.sub(r'\bshe\b', r'they', df_alter['hypothesis'].iloc[idx]) df_alter['hypothesis'].iloc[idx] = re.sub(r'\bhe\b', r'they', df_alter['hypothesis'].iloc[idx])
方法2:更高效的向量化操作(推荐)
完全不需要写循环,直接用pandas内置的字符串替换方法,性能比手动循环高很多:
df_alter['hypothesis'] = df_alter['hypothesis'].str.replace(r'\bshe\b', 'they', regex=True) df_alter['hypothesis'] = df_alter['hypothesis'].str.replace(r'\bhe\b', 'they', regex=True)
方法3:重置索引恢复连续标签
如果你希望继续用原来的写法,可以先重置行索引:
# drop=True表示丢弃原来的旧索引,不保留为新列 df_alter = df_alter.reset_index(drop=True) # 之后你原来的循环代码就可以正常运行了
内容的提问来源于stack exchange,提问作者curiousCoder
相关产品推荐
相关产品推荐

