You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用kmeans-pytorch在GPU上实现文本聚类遇内存错误(需生成超1000簇)

K-Means聚类内存错误解决方案

问题背景

使用kmeans-pytorch实现K-Means聚类时,当设置聚类数k=1000(超过10个)时触发内存错误,数据集包含7000条文本记录。

代码片段

import torch
from sklearn.feature_extraction.text import TfidfVectorizer
from kmeans_pytorch import kmeans

text_data = # list of 7000 records 

# Preprocess the data
vectorizer = TfidfVectorizer(stop_words='english')
X = vectorizer.fit_transform(text_data)

# Convert sparse matrix to PyTorch tensor
X = torch.Tensor(X.toarray())

# Move the data to the GPU
X = X.cuda()

# Run k-means clustering on the GPU
k = 1000
cluster_assignments, centroids = kmeans(X, k, device=torch.device('cuda'))

报错信息

RuntimeError: [enforce fail at alloc_cpu.cpp:73] . DefaultCPUAllocator: can't allocate memory: you tried to allocate 41359010000 bytes. Error code 12 (Cannot allocate memory)

解决方法

1. 压缩TF-IDF特征维度

TF-IDF默认生成的高维特征是内存爆炸的核心原因,可通过两种方式降维:

  • 限制特征数量:在TfidfVectorizer中设置max_features,保留最具区分度的特征,比如:
    vectorizer = TfidfVectorizer(stop_words='english', max_features=5000)
    
  • 用TruncatedSVD处理稀疏矩阵(无需转稠密):
    from sklearn.decomposition import TruncatedSVD
    svd = TruncatedSVD(n_components=300)  # 压缩到300维
    X_compressed = svd.fit_transform(X)
    X = torch.Tensor(X_compressed).cuda()
    

2. 避免稀疏矩阵转稠密Tensor

原代码中X.toarray()将稀疏矩阵转为稠密数组,直接占用大量内存,可转为PyTorch稀疏Tensor:

import numpy as np
coo = X.tocoo()
values = torch.FloatTensor(coo.data)
indices = torch.LongTensor(np.vstack((coo.row, coo.col)))
X_sparse = torch.sparse.FloatTensor(indices, values, torch.Size(coo.shape)).cuda()

注意:需确认kmeans-pytorch是否支持稀疏Tensor输入,若不支持优先选择降维方案。

3. 调整kmeans-pytorch参数优化内存

  • 设置batch_size分批计算:
    cluster_assignments, centroids = kmeans(
        X, k, device=torch.device('cuda'), batch_size=1000
    )
    
    分批计算样本与质心的距离,避免一次性加载所有数据计算距离矩阵。
  • 提前终止迭代:设置tol参数(比如tol=1e-3),当质心变化小于阈值时停止迭代,减少内存占用与计算时间。

4. 切换到CPU计算(GPU内存不足时)

若GPU内存无法承载,改用CPU并配合分批、降维策略:

cluster_assignments, centroids = kmeans(
    X, k, device=torch.device('cpu'), batch_size=1000
)

5. 清理闲置内存

运行聚类前清理PyTorch缓存与系统垃圾:

import gc
gc.collect()
torch.cuda.empty_cache()  # GPU环境下执行

内容的提问来源于stack exchange,提问作者Tanmay Shrivastava

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 20:24:32