如何使用numpy计算频谱图指定区域的最大/最小峰值?
实现方案
核心思路
利用NumPy的广播特性和布尔掩码快速筛选有效区域,全程基于NumPy底层运算实现,避免Python层循环带来的性能损耗,支持超大规模频谱数据的快速聚合计算。
两种实现版本
版本1:低内存版本(推荐大数据量使用)
通过两次一维聚合实现,内存占用极低,区间数量不大时性能最优:
import numpy as np def calculate_region_peaks( spectrogram_data: np.ndarray, freq_labels: np.ndarray, time_labels: np.ndarray, freq_bins: np.ndarray, time_bins: np.ndarray, calc_type: str = "max" ) -> np.ndarray: """ 计算频谱图指定区域的最值 参数说明: spectrogram_data: 二维频谱数据,默认shape为(频率点数, 时间点数),维度相反可提前转置 freq_labels: 一维频率标签数组,长度等于频谱数据的频率轴长度 time_labels: 一维时间标签数组,长度等于频谱数据的时间轴长度 freq_bins: 频率区间数组,shape为(M, 2),每行对应[起始频率, 结束频率] time_bins: 时间区间数组,shape为(N, 2),每行对应[起始时间, 结束时间] calc_type: 计算类型,可选"max"(最大值)或"min"(最小值) 返回值: 二维数组,shape为(M, N),对应每个频率区间×时间区间的最值 """ # 校验聚合函数 if calc_type == "max": agg_func = np.nanmax elif calc_type == "min": agg_func = np.nanmin else: raise ValueError("calc_type only support 'max' or 'min'") # 生成频率维度掩码 shape=(M, 频率点数) freq_mask = np.logical_and( freq_labels[None, :] >= freq_bins[:, [0]], freq_labels[None, :] < freq_bins[:, [1]] # 左闭右开,可按需调整边界判断 ) # 第一步:按频率区间聚合,得到中间数组 shape=(M, 时间点数) mid_arr = np.full((len(freq_bins), len(time_labels)), np.nan, dtype=spectrogram_data.dtype) for m in range(len(freq_bins)): valid_freq_data = spectrogram_data[freq_mask[m], :] if valid_freq_data.size > 0: mid_arr[m, :] = agg_func(valid_freq_data, axis=0) # 生成时间维度掩码 shape=(N, 时间点数) time_mask = np.logical_and( time_labels[None, :] >= time_bins[:, [0]], time_labels[None, :] < time_bins[:, [1]] # 左闭右开,可按需调整边界判断 ) # 第二步:按时间区间聚合,得到最终结果 shape=(M, N) res = np.full((len(freq_bins), len(time_bins)), np.nan, dtype=spectrogram_data.dtype) for n in range(len(time_bins)): valid_time_data = mid_arr[:, time_mask[n]] if valid_time_data.size > 0: res[:, n] = agg_func(valid_time_data, axis=1) return res
版本2:无循环纯广播版本(区间规模较小时更快)
完全避免Python层循环,全部运算交给NumPy底层执行,区间数量不大时速度更快,但内存占用更高:
import numpy as np def calculate_region_peaks_broadcast( spectrogram_data: np.ndarray, freq_labels: np.ndarray, time_labels: np.ndarray, freq_bins: np.ndarray, time_bins: np.ndarray, calc_type: str = "max" ) -> np.ndarray: if calc_type == "max": agg_func = np.nanmax elif calc_type == "min": agg_func = np.nanmin else: raise ValueError("calc_type only support 'max' or 'min'") F, T = spectrogram_data.shape M, N = len(freq_bins), len(time_bins) # 生成掩码 freq_mask = np.logical_and(freq_labels[None, :] >= freq_bins[:, [0]], freq_labels[None, :] < freq_bins[:, [1]]) time_mask = np.logical_and(time_labels[None, :] >= time_bins[:, [0]], time_labels[None, :] < time_bins[:, [1]]) # 维度扩展广播 spec_4d = spectrogram_data[None, :, None, :] full_mask = freq_mask[:, :, None, None] & time_mask[None, None, :, :] # 掩码替换为nan后聚合 spec_masked = np.where(full_mask, spec_4d, np.nan) return agg_func(spec_masked, axis=(1, 3))
使用示例
# 读取数据集 DATA_PATH = "你本地存储的npz文件路径" with np.load(DATA_PATH) as data: frequency_labels = data['frequency_labels'] time_labels = data['time_labels'] spectrogram_data = data['data'] # 示例输入区间 time_bins = np.array([[0.0, 10.0], [10.0, 20.0], [20.0, 30.0]]) freq_bins = np.array([[450, 550], [550, 1500], [1500, 2500]]) # 计算最大值 max_result = calculate_region_peaks(spectrogram_data, frequency_labels, time_labels, freq_bins, time_bins, calc_type="max") # 计算最小值 min_result = calculate_region_peaks(spectrogram_data, frequency_labels, time_labels, freq_bins, time_bins, calc_type="min")
注意事项
- 若你的频谱数据维度为
(时间点数, 频率点数),传入函数前先转置即可:spectrogram_data = spectrogram_data.T - 默认区间判断为左闭右开,可根据需求调整掩码中的比较运算符修改边界逻辑
- 无有效数据的区间会返回
nan,可根据业务需要后续替换为默认值
内容的提问来源于stack exchange,提问作者Flovdis
相关产品推荐
相关产品推荐

