You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

降低数组数据类型后仍触发Memory Error的原因及解决建议

欧氏距离计算内存错误问题解析与解决办法

问题背景

我有一个60660行×36列的数据集,可通过np.random.rand(60660, 36)生成同形状的随机矩阵模拟。使用sklearn.metrics.pairwise.euclidean_distances计算全量欧氏距离时,因生成60660×60660的float64数组需要27.4GiB内存,触发内存错误。尝试将输入矩阵转为float16后,仍报相同的内存错误,且测试小数据集时发现输出始终是float64类型。

复现代码与报错

初始代码(float64输入):

import numpy as np
from sklearn.metrics.pairwise import euclidean_distances

X = np.random.rand(60660, 36)
distances = euclidean_distances(X, squared=True)

报错:

MemoryError: Unable to allocate 27.4 GiB for an array with shape (60660, 60660) and data type float64

float16输入代码:

X = np.random.rand(60660, 36).astype('float16')
distances = euclidean_distances(X, squared=True)

报错核心片段:

MemoryError: Unable to allocate 27.4 GiB for an array with shape (60660, 60660) and data type float64

原因分析

sklearn的euclidean_distances函数为避免计算过程中的精度损失,会自动将输入数据向上转换为float64执行矩阵乘法等核心运算,最终输出的距离矩阵固定为float64类型。因此即使输入是float16/float32,内存需求依然由float64的全量矩阵决定,无法通过降低输入类型减少内存占用。

从报错堆栈可见,核心运算-2 * safe_sparse_dot(X, Y.T)会强制使用float64计算,这是导致内存问题的根本原因。

解决办法

1. 分块计算(内存友好的精确方案)

不一次性生成完整距离矩阵,而是分批次计算部分行的距离,按需存储或处理结果,大幅降低瞬时内存占用。

import numpy as np
from sklearn.metrics.pairwise import euclidean_distances

X = np.random.rand(60660, 36).astype('float32')  # 用float32减少输入内存占用
chunk_size = 1000  # 可根据自身内存容量调整
total_chunks = (X.shape[0] + chunk_size - 1) // chunk_size

# 逐块计算并处理结果(示例为写入磁盘,可替换为业务逻辑)
for i in range(total_chunks):
    start = i * chunk_size
    end = min((i+1)*chunk_size, X.shape[0])
    chunk_dist = euclidean_distances(X[start:end], X, squared=True).astype('float32')
    np.savez(f"distance_chunk_{i}.npz", dist=chunk_dist)  # 写入磁盘释放内存
    print(f"完成第{i+1}/{total_chunks}块计算")

说明:每次仅占用chunk_size × 60660 × 4字节(float32)的内存,比如chunk_size=1000时仅约232MB,完全适配普通机器内存。

2. 近似最近邻算法(非全量距离场景)

若业务仅需每个样本的K近邻距离而非全量矩阵,可使用近似最近邻库(如FAISS),无需生成完整距离矩阵,内存占用极低。

import faiss
import numpy as np

X = np.random.rand(60660, 36).astype('float32')
index = faiss.IndexFlatL2(36)  # L2距离等价于欧氏距离平方
index.add(X)
k = 10  # 设定每个样本需获取的近邻数量
distances, indices = index.search(X, k)  # 输出为(60660, k)的float32矩阵

说明:仅存储每个样本的K个近邻距离,内存占用仅约2.3MB(60660×10×4字节),计算效率远高于全量距离计算。

3. 手动实现低精度全量计算(精度换内存)

若必须生成完整距离矩阵且能接受精度损失,可手动基于欧氏距离的矩阵公式实现计算,强制使用低精度类型存储结果:

import numpy as np

X = np.random.rand(60660, 36).astype('float32')
XX = np.sum(X**2, axis=1, keepdims=True)
# 基于公式:dist(a,b)² = ||a||² + ||b||² - 2a·b
distances = XX + XX.T - 2 * np.dot(X, X.T)
distances = np.maximum(distances, 0)  # 修正浮点误差导致的负数

# 可选:进一步转为float16减少内存(精度损失更大)
# distances = distances.astype('float16')

说明:float32版本内存占用约13.7GiB(为float64的一半),float16版本约6.8GiB,但需注意float16的精度可能无法满足业务需求,需提前验证。


内容的提问来源于stack exchange,提问作者Mike Jadwin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 14:41:32