Pandas遍历Series遇KeyError:如何检查NumPy浮点值是否大于0?
问题描述
我有一个名为df1的DataFrame,结构如下:
Open Close DiffMa %Percentage Open Close DiffMa2 %Percentage2 2022-11-04 13:30:00-04:00 42.099998 42.224998 -0.135001 0.296912 13.9586 14.0150 -0.06584 0.404054 2022-11-04 14:30:00-04:00 42.220001 42.330002 0.028002 0.260541 14.0150 14.0550 0.01816 0.285408 2022-11-04 15:30:00-04:00 42.320000 42.630001 0.311001 0.732517 14.0600 14.1100 0.07716 0.355613 2022-11-07 09:30:00-05:00 43.049999 42.294998 -0.019002 -1.753777 14.3200 14.0299 -0.00308 -2.025839 2022-11-07 10:30:00-05:00 42.299999 42.195000 -0.140000 -0.248226 14.0300 13.9865 -0.05278 -0.310050
df1['%Percentage'][0]的输出为:0.2969121247755807
我需要统计%Percentage和%Percentage2两列同时为正或同时为负的次数,但编写循环时触发KeyError,简化后的循环代码:
countera = 0 counterb = 0 for i in df1['%Percentage']: if df1['%Percentage'][i] > 0: countera = 0
报错信息:
--------------------------------------------------------------------------- KeyError Traceback (most recent call last) <ipython-input-10-cc9f6f8fa7d5> in <module> 4 5 for i in df1['%Percentage']: ----> 6 if df1['%Percentage'][i] > 0: 7 countera += 1 8 ~/opt/anaconda3/envs/StockPredictionGameification/lib/python3.6/site-packages/pandas/core/series.py in __getitem__(self, key) 880 881 elif key_is_scalar: --> 882 return self._get_value(key) 883 884 if is_hashable(key): ~/opt/anaconda3/envs/StockPredictionGameification/lib/python3.6/site-packages/pandas/core/series.py in _get_value(self, label, takeable) 988 989 # Similar to Index.get_value, but we do not fall back to positional --> 990 loc = self.index.get_loc(label) 991 return self.index._get_values_for_loc(self, loc, label) 992 ~/opt/anaconda3/envs/StockPredictionGameification/lib/python3.6/site-packages/pandas/core/indexes/datetimes.py in get_loc(self, key, method, tolerance) 620 else: 621 # unrecognized type --> 622 raise KeyError(key) 623 624 try: KeyError: 0.2969121247755807
我猜测是浮点数精度问题,但不确定具体原因,需要问题解析和解决办法。
问题原因
你遇到的KeyError和浮点数精度无关,核心问题是循环变量的理解错误:
for i in df1['%Percentage']遍历的是该列的数值(比如第一个i是0.2969121247755807)- 而
df1['%Percentage'][i]是试图用这个数值作为索引去查找数据,但你的DataFrame索引是时间戳(比如2022-11-04 13:30:00-04:00),时间戳索引里不存在0.2969121247755807这个值,因此触发KeyError。
解决方法
方法1:修复循环逻辑
如果一定要用循环,可以遍历索引或同时遍历两列的数值对:
方式A:遍历索引
counter_both_pos = 0 counter_both_neg = 0 for idx in df1.index: p1 = df1['%Percentage'][idx] p2 = df1['%Percentage2'][idx] if p1 > 0 and p2 > 0: counter_both_pos += 1 elif p1 < 0 and p2 < 0: counter_both_neg += 1 print(f"同时为正次数:{counter_both_pos}") print(f"同时为负次数:{counter_both_neg}")
方式B:直接遍历两列的数值对
counter_both_pos = 0 counter_both_neg = 0 for p1, p2 in zip(df1['%Percentage'], df1['%Percentage2']): if p1 > 0 and p2 > 0: counter_both_pos += 1 elif p1 < 0 and p2 < 0: counter_both_neg += 1 print(f"同时为正次数:{counter_both_pos}") print(f"同时为负次数:{counter_both_neg}")
方法2:Pandas向量化操作(推荐)
Pandas的核心优势是向量化运算,比循环高效得多,无需遍历每一行:
# 计算同时为正的布尔序列,sum()统计True的数量 both_pos = ((df1['%Percentage'] > 0) & (df1['%Percentage2'] > 0)).sum() # 计算同时为负的布尔序列 both_neg = ((df1['%Percentage'] < 0) & (df1['%Percentage2'] < 0)).sum() print(f"同时为正次数:{both_pos}") print(f"同时为负次数:{both_neg}")
这种方式代码更简洁,执行效率远高于循环,尤其当DataFrame数据量较大时差距明显。
内容的提问来源于stack exchange,提问作者Niko
相关产品推荐
相关产品推荐

