You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在pandas DataFrame指定范围内应用掩码计算过载与基线均值

指定区间过载/基线均值计算实现方案

核心修正思路

原有代码存在两处问题导致结果不符合预期:

  1. 语法错误:np.where(mask == regression area.index) 存在空格变量名问题,且直接将numpy数组和索引对象做相等匹配,无法得到正确的筛选结果
  2. 逻辑错误:直接对mask数组做[first:last]切片无法得到目标区间内的有效过载索引,会导致索引对齐错误

修正逻辑如下:

  • 保留原有mask生成逻辑,支持no_overload_cycles动态调整
  • 提取[first, last)区间的索引范围后,分别筛选出属于mask、不属于mask的两类索引
  • 基于筛选后的索引直接取值计算均值,避免索引对齐问题
  • 将mask转为集合结构,大幅提升索引成员判断的效率

完整可运行代码

import pandas as pd
import numpy as np

# 可动态调整的参数
mask_number = 5
no_overload_cycles = 1
first = 4
last = 17

# 测试数据
hyst = pd.DataFrame({"test":[12, 4, 5, 4, 1, 3, 2, 5, 10, 9, 7, 5, 3, 6, 3, 2 ,1, 5, 2]})

# 生成原始过载掩码(和原有逻辑完全一致)
mask_list = []
for i in range(0, len(hyst)-1, mask_number):
    for x in range(no_overload_cycles):
        mask_list.append(i+x)
mask = np.array(mask_list)
mask_set = set(mask)

# 提取目标区间
regression_index = hyst.iloc[first:last].index
# 筛选区间内的过载索引、基线索引
overload_idx = [idx for idx in regression_index if idx in mask_set]
baseline_idx = [idx for idx in regression_index if idx not in mask_set]

# 计算均值
mean_range_overload = hyst.loc[overload_idx, 'test'].mean()
mean_range_baseline = hyst.loc[baseline_idx, 'test'].mean()

结果验证

场景1:no_overload_cycles=1

运行后输出完全匹配预期:

>>> print(mean_range_overload)
4.0
>>> print(mean_range_baseline)
4.545454545454546

其中过载点对应索引5、10、15,取值为3、7、2,均值计算正确;基线点为区间内排除上述三个索引的剩余10个值,均值符合预期。

场景2:no_overload_cycles=3

此时生成的mask为[ 0 1 2 5 6 7 10 11 12 15 16 17 20],落在目标区间内的过载索引为5、6、7、10、11、12、15、16,计算得到:

  • 过载均值:3.5
  • 基线均值:5.8
    完全符合逻辑要求。

内容的提问来源于stack exchange,提问作者Murray Ross

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 05:01:11