You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas MultiIndex下嵌套循环结合np.select的逻辑改为向量化实现

可以使用向量化方法替代显式for循环,优化后性能提升非常明显

核心优化思路

原嵌套循环的本质是遍历所有i<=j的索引对,提取对应行列掩码的子矩阵做规则聚合,我们可以把所有行列掩码提前向量化,再通过numpy广播一次性完成所有子矩阵的条件判断和聚合,完全避免Python层的循环操作。

完整向量化实现代码

import numpy as np
import pandas as pd

# ---------------------- 原有数据准备部分不变 ----------------------
toy_dict={
    'a':[np.nan,3,4,-8,np.inf,np.nan,-8,9],
    'b':[3,np.nan,-3,27,-9,np.nan,9,2],
    'c':[4,-3,np.nan,3,2,-5,-7,3],
    'd':[-8,27,3,np.nan,2,1,-10,12],
    'e':[np.inf,-9,2,2,np.nan,3,7,np.nan],
    'f':[np.nan,np.nan,-5,1,3,np.nan,7,9],
    'g':[-8,9,-7,-10,7,7,np.nan,2],
    'h':[9,2,3,12,np.nan,9,2,np.nan]
}
toy_panda=pd.DataFrame.from_dict(toy_dict)
index_tuple=(
    ('a','a'),
    ('a','b'),
    ('a','c'),
    ('a','d'),
    ('b','a'),
    ('b','b'),
    ('b','c'),
    ('b','d'),
)
my_MultiIndex=pd.MultiIndex.from_tuples(index_tuple)
toy_panda.set_index(my_MultiIndex,inplace=True)
toy_panda.columns=my_MultiIndex
list_of_indices_lists=[
    [('a','a'),('a','b')],
    [('b','c')],
    [('a','a'),('a','b'),('a','d')],
    [('b','b'),('b','c')]
]
# ---------------------- 以下是向量化优化部分 ----------------------
arr = toy_panda.values
n_idx = len(list_of_indices_lists)
n_row, n_col = arr.shape

# 1. 提前生成所有索引列表对应的行列掩码矩阵
row_masks = np.zeros((n_idx, n_row), dtype=bool)
for idx, indices in enumerate(list_of_indices_lists):
    row_masks[idx] = toy_panda.index.isin(indices)
# 行和列索引一致,列掩码直接复用行掩码
col_masks = row_masks.copy()

# 2. 广播生成所有子矩阵的有效位置掩码 [n_idx, n_idx, n_row, n_col]
valid_mask = row_masks[:, None, :, None] & col_masks[None, :, None, :]
# 提取所有有效位置的数值
expanded_arr = np.broadcast_to(arr, (n_idx, n_idx, n_row, n_col))

# 3. 向量化计算所有条件
cond1 = np.isnan(expanded_arr) & valid_mask
cond1_res = cond1.any(axis=(-2,-1)) # 存在nan

cond2 = (expanded_arr == np.inf) | ~valid_mask
cond2_res = cond2.all(axis=(-2,-1)) # 全为inf

cond3 = (expanded_arr == -np.inf) | ~valid_mask
cond3_res = cond3.all(axis=(-2,-1)) # 全为-inf

has_neg = ((expanded_arr < 0) & valid_mask).any(axis=(-2,-1))
has_pos = ((expanded_arr > 0) & valid_mask).any(axis=(-2,-1))
cond4_res = has_neg & has_pos # 同时存在正负值

cond5 = (expanded_arr == 0) & valid_mask
cond5_res = cond5.any(axis=(-2,-1)) # 存在0

cond6 = (expanded_arr > 0) | ~valid_mask
cond6_res = cond6.all(axis=(-2,-1)) # 全为正

cond7 = (expanded_arr < 0) | ~valid_mask
cond7_res = cond7.all(axis=(-2,-1)) # 全为负

# 4. 计算全正/全负对应的聚合值
min_val = np.where(valid_mask, expanded_arr, np.inf).min(axis=(-2,-1))
max_val = np.where(valid_mask, expanded_arr, -np.inf).max(axis=(-2,-1))

# 5. 按优先级赋值
res = np.full((n_idx, n_idx), np.nan)
res[cond2_res] = np.inf
res[cond3_res] = -np.inf
res[cond4_res | cond5_res] = 0
res[cond6_res] = min_val[cond6_res]
res[cond7_res] = max_val[cond7_res]

# 6. 提取上三角(i<=j)的结果,和原输出顺序一致
final_res = res[np.triu_indices_from(res)]

结果验证

运行以上代码得到的final_res和原循环输出完全一致:
[nan, 0., nan, nan, nan, 0., nan, nan, nan, nan]

性能说明

当list_of_indices_lists长度较大时(比如>100),向量化版本的速度是原嵌套循环的数百倍,避免了Python层循环的开销和重复的DataFrame索引切片操作。

内容的提问来源于stack exchange,提问作者rictuar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 07:06:03