You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中转换为np.array时内存饱和及内核重启问题求助

解决批量转NumPy数组内存饱和的方法

原代码里np.array(processed_X)会触发内存翻倍——列表里的每个子数组已经占了内存,转成大数组时又要重新分配一份完整内存空间,两份数据同时存在直接撑爆内存。以下是针对性的解决办法:

1. 预先分配内存,直接填充数据

提前计算好最终数组的形状,初始化一个空数组,循环时直接把处理好的样本填进去,避免列表存储子数组的额外开销,也不会出现两份数据占内存的情况。

修改后的完整代码:

import numpy as np
import scipy.signal
from skimage.transform import resize
from tqdm import tqdm

def apply_stft(signal, nperseg=256, noverlap=128):
    f, t, Zxx = scipy.signal.stft(signal, nperseg=nperseg, noverlap=noverlap)
    return np.abs(Zxx)

def resize_stft(stft_result, output_shape=(224, 224)):
    return resize(stft_result, output_shape, mode='constant')

def process_dataset(X):
    # 先获取单个样本处理后的形状,用来预分配大数组
    first_example = X[0]
    stft_results = [apply_stft(first_example[:, i]) for i in range(first_example.shape[1])]
    resized_stft = [resize_stft(stft) for stft in stft_results]
    single_sample_shape = np.stack(resized_stft, axis=-1).shape
    # 清理临时变量
    del stft_results, resized_stft, first_example

    # 预分配最终数组,用float32替代默认float64,直接省一半内存
    final_dtype = np.float32
    processed_X = np.zeros((len(X),) + single_sample_shape, dtype=final_dtype)

    for idx, example in tqdm(enumerate(X), total=len(X)):
        stft_results = [apply_stft(example[:, i]) for i in range(example.shape[1])]
        resized_stft = [resize_stft(stft) for stft in stft_results]
        stacked_stft = np.stack(resized_stft, axis=-1).astype(final_dtype)
        processed_X[idx] = stacked_stft
        # 及时释放临时变量
        del stft_results, resized_stft, stacked_stft
    print(len(processed_X))
    print('FINISHED !!!')
    return processed_X

2. 降低数据精度

如果任务对精度要求不高,把数据从默认的float64转成float32甚至float16(需确认下游模型支持),内存占用直接减半或降到1/4,这是最有效的内存压缩手段之一。

3. 分批次处理并缓存到磁盘

如果预分配内存还是不够,就分小批次处理,每处理完一批就保存到磁盘,不用把所有数据都放在内存里。后续训练或分析时再按需加载批次数据。

示例代码片段:

import os
import numpy as np
import scipy.signal
from skimage.transform import resize
from tqdm import tqdm

# 复用前面的apply_stft和resize_stft函数

def process_dataset_batch(X, batch_size=64, save_dir='processed_batches'):
    os.makedirs(save_dir, exist_ok=True)
    total_batches = len(X) // batch_size + (1 if len(X) % batch_size != 0 else 0)
    for batch_idx in tqdm(range(total_batches)):
        start_idx = batch_idx * batch_size
        end_idx = min((batch_idx+1)*batch_size, len(X))
        batch_data = X[start_idx:end_idx]
        
        processed_batch = []
        for example in batch_data:
            stft_results = [apply_stft(example[:, i]) for i in range(example.shape[1])]
            resized_stft = [resize_stft(stft) for stft in stft_results]
            stacked_stft = np.stack(resized_stft, axis=-1).astype(np.float32)
            processed_batch.append(stacked_stft)
            del stft_results, resized_stft, stacked_stft
        
        # 保存批次到磁盘,用npz格式压缩存储
        np.savez(f'{save_dir}/batch_{batch_idx}.npz', data=np.array(processed_batch))
        # 清理当前批次的内存
        del processed_batch, batch_data
    print('FINISHED !!!')

4. 强制触发垃圾回收

有时候Python的垃圾回收不会立即清理已删除的变量,可以在循环末尾手动调用gc.collect()强制释放内存:

import gc
# 在循环内的del语句后添加
gc.collect()

内容的提问来源于stack exchange,提问作者houfis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 02:06:04