使用.loc遍历Dataframe时触发KeyError的问题排查
问题描述
我用DataFrame存储了包含index、name、data列的数据,并保存为pkl文件。在Jupyter Notebook中逐行使用df.loc可以正常读取数据,但编写脚本通过for循环遍历DataFrame时触发KeyError,尝试.loc和.iloc都无法解决。
相关代码
str_comp= 'something' df= pd.read_pickle(<filepath>) data_to_save= None for i in df: if df.loc[i,'name'] == str_comp: data_to_save= df.loc[i,'data'] continue
报错信息
415 raise KeyError(key) from err 416 if isinstance(key, Hashable): --> 417 raise KeyError(key) 418 self._check_indexing_error(key) 419 raise KeyError(key) KeyError: 'data'
完整报错栈
KeyError Traceback (most recent call last) Cell In[28], line 6 4 if var in dict: 5 if var1 != 'str': ----> 6 save_to_data(var,"var2") File ~<filepath to script>/script.py:48, in save_to_data(name, y) 47 def save_to_data(name,y): ----> 48 data= df_creation(name,y) 49 data_to_pickle(f"<filepath to data folder/<name of file>.pkl") File <filepath to script>\script.py:31, in df_creation(name, y) data_to_save= None for i in df: if df.loc[i,'name'] == str_comp: data_to_save= df.loc[i,'data'] continue File<filepath to directory>\venv\Lib\site-packages\pandas\core\indexing.py:1183, in _LocationIndexer.__getitem__(self, key) 1181 key = tuple(com.apply_if_callable(x, self.obj) for x in key) 1182 if self._is_scalar_access(key): -> 1183 return self.obj._get_value(*key, takeable=self._takeable) 1184 return self._getitem_tuple(key) 1185 else: 1186 # we by definition only have the 0th axis File <filepath to directory>\venv\Lib\site-packages\pandas\core\frame.py:4209, in DataFrame._get_value(self, index, col, takeable) 4203 engine = self.index._engine 4205 if not isinstance(self.index, MultiIndex): 4206 # CategoricalIndex: Trying to use the engine fastpath may give incorrect 4207 # results if our categories are integers that dont match our codes 4208 # IntervalIndex: IntervalTree has no get_loc -> 4209 row = self.index.get_loc(index) 4210 return series._values[row] 4212 # For MultiIndex going through engine effectively restricts us to 4213 # same-length tuples; see test_get_set_value_no_partial_indexing File<filepath to directory>\venv\Lib\site-packages\pandas\core\indexes\range.py:417, in RangeIndex.get_loc(self, key) 415 raise KeyError(key) from err 416 if isinstance(key, Hashable): -> 417 raise KeyError(key) 418 self._check_indexing_error(key) 419 raise KeyError(key) KeyError: 'name'
解决方案
问题根源
for i in df这种遍历方式,本质是遍历DataFrame的列名,而非行索引。当循环变量i取到name、data这类列名时,df.loc[i, 'name']会把列名当成行索引去查找,而你的行索引里没有这些字符串值,自然触发KeyError。
修复方案1:遍历行索引
把循环改成遍历DataFrame的行索引集合,这样i就是每行的索引值,再用.loc就能正常定位行和列:
str_comp= 'something' df= pd.read_pickle(<filepath>) data_to_save= None # 遍历行索引而非列名 for i in df.index: if df.loc[i,'name'] == str_comp: data_to_save= df.loc[i,'data'] break # 找到目标后直接退出循环,没必要继续遍历
修复方案2:用布尔索引直接筛选(更高效)
pandas不推荐用循环遍历行,更高效的方式是直接用布尔索引筛选匹配的行,一次性获取目标数据:
str_comp= 'something' df= pd.read_pickle(<filepath>) # 筛选name等于str_comp的行,提取data列 matched_data = df.loc[df['name'] == str_comp, 'data'] # 如果确定只有一个匹配项,直接取第一个值 if not matched_data.empty: data_to_save = matched_data.iloc[0]
为什么Jupyter里能正常运行?
因为你在Jupyter里是手动指定行索引(比如df.loc[0, 'name']),而脚本里的循环逻辑错误地遍历了列名,两者的索引定位逻辑完全不同。
内容的提问来源于stack exchange,提问作者Desmond Spicer
相关产品推荐
相关产品推荐

