Python中转换为np.array时内存饱和及内核重启问题求助
解决批量转NumPy数组内存饱和的方法
原代码里np.array(processed_X)会触发内存翻倍——列表里的每个子数组已经占了内存,转成大数组时又要重新分配一份完整内存空间,两份数据同时存在直接撑爆内存。以下是针对性的解决办法:
1. 预先分配内存,直接填充数据
提前计算好最终数组的形状,初始化一个空数组,循环时直接把处理好的样本填进去,避免列表存储子数组的额外开销,也不会出现两份数据占内存的情况。
修改后的完整代码:
import numpy as np import scipy.signal from skimage.transform import resize from tqdm import tqdm def apply_stft(signal, nperseg=256, noverlap=128): f, t, Zxx = scipy.signal.stft(signal, nperseg=nperseg, noverlap=noverlap) return np.abs(Zxx) def resize_stft(stft_result, output_shape=(224, 224)): return resize(stft_result, output_shape, mode='constant') def process_dataset(X): # 先获取单个样本处理后的形状,用来预分配大数组 first_example = X[0] stft_results = [apply_stft(first_example[:, i]) for i in range(first_example.shape[1])] resized_stft = [resize_stft(stft) for stft in stft_results] single_sample_shape = np.stack(resized_stft, axis=-1).shape # 清理临时变量 del stft_results, resized_stft, first_example # 预分配最终数组,用float32替代默认float64,直接省一半内存 final_dtype = np.float32 processed_X = np.zeros((len(X),) + single_sample_shape, dtype=final_dtype) for idx, example in tqdm(enumerate(X), total=len(X)): stft_results = [apply_stft(example[:, i]) for i in range(example.shape[1])] resized_stft = [resize_stft(stft) for stft in stft_results] stacked_stft = np.stack(resized_stft, axis=-1).astype(final_dtype) processed_X[idx] = stacked_stft # 及时释放临时变量 del stft_results, resized_stft, stacked_stft print(len(processed_X)) print('FINISHED !!!') return processed_X
2. 降低数据精度
如果任务对精度要求不高,把数据从默认的float64转成float32甚至float16(需确认下游模型支持),内存占用直接减半或降到1/4,这是最有效的内存压缩手段之一。
3. 分批次处理并缓存到磁盘
如果预分配内存还是不够,就分小批次处理,每处理完一批就保存到磁盘,不用把所有数据都放在内存里。后续训练或分析时再按需加载批次数据。
示例代码片段:
import os import numpy as np import scipy.signal from skimage.transform import resize from tqdm import tqdm # 复用前面的apply_stft和resize_stft函数 def process_dataset_batch(X, batch_size=64, save_dir='processed_batches'): os.makedirs(save_dir, exist_ok=True) total_batches = len(X) // batch_size + (1 if len(X) % batch_size != 0 else 0) for batch_idx in tqdm(range(total_batches)): start_idx = batch_idx * batch_size end_idx = min((batch_idx+1)*batch_size, len(X)) batch_data = X[start_idx:end_idx] processed_batch = [] for example in batch_data: stft_results = [apply_stft(example[:, i]) for i in range(example.shape[1])] resized_stft = [resize_stft(stft) for stft in stft_results] stacked_stft = np.stack(resized_stft, axis=-1).astype(np.float32) processed_batch.append(stacked_stft) del stft_results, resized_stft, stacked_stft # 保存批次到磁盘,用npz格式压缩存储 np.savez(f'{save_dir}/batch_{batch_idx}.npz', data=np.array(processed_batch)) # 清理当前批次的内存 del processed_batch, batch_data print('FINISHED !!!')
4. 强制触发垃圾回收
有时候Python的垃圾回收不会立即清理已删除的变量,可以在循环末尾手动调用gc.collect()强制释放内存:
import gc # 在循环内的del语句后添加 gc.collect()
内容的提问来源于stack exchange,提问作者houfis
相关产品推荐
相关产品推荐

