降低数组数据类型后仍触发Memory Error的原因及解决建议
问题背景
我有一个60660行×36列的数据集,可通过np.random.rand(60660, 36)生成同形状的随机矩阵模拟。使用sklearn.metrics.pairwise.euclidean_distances计算全量欧氏距离时,因生成60660×60660的float64数组需要27.4GiB内存,触发内存错误。尝试将输入矩阵转为float16后,仍报相同的内存错误,且测试小数据集时发现输出始终是float64类型。
复现代码与报错
初始代码(float64输入):
import numpy as np from sklearn.metrics.pairwise import euclidean_distances X = np.random.rand(60660, 36) distances = euclidean_distances(X, squared=True)
报错:
MemoryError: Unable to allocate 27.4 GiB for an array with shape (60660, 60660) and data type float64
float16输入代码:
X = np.random.rand(60660, 36).astype('float16') distances = euclidean_distances(X, squared=True)
报错核心片段:
MemoryError: Unable to allocate 27.4 GiB for an array with shape (60660, 60660) and data type float64
原因分析
sklearn的euclidean_distances函数为避免计算过程中的精度损失,会自动将输入数据向上转换为float64执行矩阵乘法等核心运算,最终输出的距离矩阵固定为float64类型。因此即使输入是float16/float32,内存需求依然由float64的全量矩阵决定,无法通过降低输入类型减少内存占用。
从报错堆栈可见,核心运算-2 * safe_sparse_dot(X, Y.T)会强制使用float64计算,这是导致内存问题的根本原因。
解决办法
1. 分块计算(内存友好的精确方案)
不一次性生成完整距离矩阵,而是分批次计算部分行的距离,按需存储或处理结果,大幅降低瞬时内存占用。
import numpy as np from sklearn.metrics.pairwise import euclidean_distances X = np.random.rand(60660, 36).astype('float32') # 用float32减少输入内存占用 chunk_size = 1000 # 可根据自身内存容量调整 total_chunks = (X.shape[0] + chunk_size - 1) // chunk_size # 逐块计算并处理结果(示例为写入磁盘,可替换为业务逻辑) for i in range(total_chunks): start = i * chunk_size end = min((i+1)*chunk_size, X.shape[0]) chunk_dist = euclidean_distances(X[start:end], X, squared=True).astype('float32') np.savez(f"distance_chunk_{i}.npz", dist=chunk_dist) # 写入磁盘释放内存 print(f"完成第{i+1}/{total_chunks}块计算")
说明:每次仅占用chunk_size × 60660 × 4字节(float32)的内存,比如chunk_size=1000时仅约232MB,完全适配普通机器内存。
2. 近似最近邻算法(非全量距离场景)
若业务仅需每个样本的K近邻距离而非全量矩阵,可使用近似最近邻库(如FAISS),无需生成完整距离矩阵,内存占用极低。
import faiss import numpy as np X = np.random.rand(60660, 36).astype('float32') index = faiss.IndexFlatL2(36) # L2距离等价于欧氏距离平方 index.add(X) k = 10 # 设定每个样本需获取的近邻数量 distances, indices = index.search(X, k) # 输出为(60660, k)的float32矩阵
说明:仅存储每个样本的K个近邻距离,内存占用仅约2.3MB(60660×10×4字节),计算效率远高于全量距离计算。
3. 手动实现低精度全量计算(精度换内存)
若必须生成完整距离矩阵且能接受精度损失,可手动基于欧氏距离的矩阵公式实现计算,强制使用低精度类型存储结果:
import numpy as np X = np.random.rand(60660, 36).astype('float32') XX = np.sum(X**2, axis=1, keepdims=True) # 基于公式:dist(a,b)² = ||a||² + ||b||² - 2a·b distances = XX + XX.T - 2 * np.dot(X, X.T) distances = np.maximum(distances, 0) # 修正浮点误差导致的负数 # 可选:进一步转为float16减少内存(精度损失更大) # distances = distances.astype('float16')
说明:float32版本内存占用约13.7GiB(为float64的一半),float16版本约6.8GiB,但需注意float16的精度可能无法满足业务需求,需提前验证。
内容的提问来源于stack exchange,提问作者Mike Jadwin

