如何对(858,891)的NumPy数组按每4列取均值生成(858,223)数组?
问题:按固定列数计算均值并生成新数组
我有一个形状为(858, 891)的NumPy数组:
[[0.41128411 0.41473465 0.41853683 ... 0.49271687 0.49322761 0.49389444] [0.40997746 0.41332964 0.41707197 ... 0.49285828 0.49338905 0.49406584] [0.40780573 0.41098776 0.41461104 ... 0.49297105 0.49353365 0.49422911] ... [0.62126761 0.61681614 0.61088054 ... 0.66176147 0.66263369 0.66313708] [0.62241648 0.61767095 0.61128526 ... 0.66124009 0.66208772 0.6625708 ] [0.62377412 0.61880136 0.6120562 ... 0.66079591 0.66163067 0.66210466]]
需要按每4列计算均值,将后续每4列的均值堆叠起来,最终得到形状为(858,223)的新数组(最后一次迭代处理3列)。目前手动处理前8列的代码如下:
a = arr[:, 0:4] b = arr[:, 4:8] a = a.mean(axis=1) b = b.mean(axis=1) new = np.column_stack((a,b))
得到的结果示例:
array([[0.41670609, 0.42968413], [0.41529677, 0.42850663], [0.41293143, 0.42641906], ..., [0.61311844, 0.57922079], [0.6136733 , 0.57750982], [0.61455882, 0.57644654]])
请问如何通过循环遍历原数组实现上述需求?
解决方案
方法一:循环遍历实现
直接遍历列的起始索引,每次截取对应列区间计算均值,再逐步堆叠结果:
import numpy as np # 假设arr为原始输入数组 result_list = [] total_cols = arr.shape[1] step = 4 for start in range(0, total_cols, step): # 确定当前切片的结束位置,防止越界 end = min(start + step, total_cols) # 计算当前列区间的行均值 current_mean = arr[:, start:end].mean(axis=1) # 将均值结果加入列表 result_list.append(current_mean) # 将列表中的数组按列堆叠成最终数组 new_arr = np.column_stack(result_list)
由于891 = 4*222 + 3,循环到最后一次时,start=888,end=min(888+4, 891)=891,会自动截取最后3列计算均值,最终得到形状为(858,223)的数组,完全符合需求。
方法二:NumPy向量化操作(推荐)
如果不需要强制使用循环,用NumPy内置的数组拆分方法可以更高效完成任务,避免循环的性能开销:
import numpy as np # 将原数组按每4列一组拆分,最后一组自动保留剩余的3列 col_groups = np.array_split(arr, np.arange(step, arr.shape[1], step), axis=1) # 对每组计算行均值,再按列堆叠 new_arr = np.column_stack([group.mean(axis=1) for group in col_groups])
np.array_split会自动处理不足4列的最后一组,结果和循环方法一致,但执行效率更高,适合处理大规模数组。
内容的提问来源于stack exchange,提问作者Vladimir Sotskov
相关产品推荐
相关产品推荐

