You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于numpy数组掩码计算均值时~取反触发KeyError的解决方法

报错原因
  • 你当前生成的mask是存储目标行位置的整数数组,不是布尔型掩码:打印输出的[0 1 5 6 10 11 15 16]是需要选中的行序号,不是逐行标记是否选中的布尔值序列。
  • 波浪符~的逻辑取反效果仅对布尔类型数组生效,对整数数组使用~会执行按位取反运算,最终得到[-1,-2,-6...]这类不存在的负数值。
  • 你调用的.loc是按索引标签筛选行的索引器,传入这些不存在的负索引时,自然会触发找不到对应行的KeyError。
修正方案

优先使用布尔掩码实现,逻辑最清晰,也不会因为DataFrame索引变化出现筛选错误:

import pandas as pd
import numpy as np

mask_number = 5
no_overload_cycles = 2
hyst = pd.DataFrame({"test":[12, 4, 5, 4, 1, 3, 2, 5, 10, 9, 7, 5, 3, 6, 3, 2 ,1, 5, 2]})

list_test = []
for i in range(0,len(hyst)-1,mask_number):
    for x in range(no_overload_cycles):
        list_test.append(i+x)
    
# 构造和DataFrame行数等长的布尔掩码,初始全为未选中(False)
mask = np.zeros(len(hyst), dtype=bool)
# 把目标位置标记为选中(True)
mask[list_test] = True

overload_mean = hyst.loc[mask, 'test'].mean() 
baseline_mean = hyst.loc[~mask, 'test'].mean()

运行后计算结果:

  • overload_mean(掩码覆盖位置均值):4.75
  • baseline_mean(掩码外位置均值):4.8

如果你想保留原来的整数位置列表写法,需要改用按位置筛选的索引器.iloc,且不能直接对整数数组用~取反,要单独生成未被覆盖的位置数组:

all_pos = np.arange(len(hyst))
# 取全量位置和已选位置的差集,就是掩码外的位置
non_mask_pos = np.setdiff1d(all_pos, list_test)

overload_mean = hyst.iloc[list_test, hyst.columns.get_loc('test')].mean()
baseline_mean = hyst.iloc[non_mask_pos, hyst.columns.get_loc('test')].mean()

注意:不要混淆.loc和.iloc的用法:.loc匹配行的索引标签,.iloc匹配行的整数位置,当DataFrame索引不是默认从0开始的连续整数时,直接传位置整数给.loc会出现筛选错误。

内容的提问来源于stack exchange,提问作者Murray Ross

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 04:27:20