You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用CamemBERT提取特征后用KNN分类时出现内存不足RuntimeError

CamemBERT特征提取+KNN分类内存不足问题解决

问题场景

我尝试用CamemBERT模型提取文本特征,再用KNN分类器对特征向量分类,编写的代码如下:

import torch
from transformers import AutoTokenizer, CamembertModel
from sklearn.neighbors import KNeighborsClassifier

tokenizer = AutoTokenizer.from_pretrained("camembert-base")
model = CamembertModel.from_pretrained("camembert-base")

data = df.to_dict(orient='split')
data = dict(zip(data['index'], data['data']))

# Collect all the input texts into a list of strings
input_texts = [str(text) for text in data.values()]

# Tokenize all the input texts together
inputs = tokenizer(input_texts, return_tensors="pt", padding=True, truncation=True)

# Get the model outputs for all the input texts
with torch.no_grad():
    outputs = model(**inputs)

# Extract the last hidden states and convert them to a numpy array
last_hidden_states = outputs.last_hidden_state
input_features = last_hidden_states[:, 0, :].numpy()

# Extract the labels from the data dictionary
input_labels = list(data.keys())

neigh = KNeighborsClassifier(n_neighbors=3)
neigh.fit(input_features, input_labels)

报错信息

运行代码后出现内存不足错误:

RuntimeError: [enforce fail at ..\c10\core\impl\alloc_cpu.cpp:72] data. DefaultCPUAllocator: not enough memory: you tried to allocate 19209424896 bytes.

数据格式

使用的数据字典格式如下:

{
    'index': [row_index_1, row_index_2, ...],
    'columns': [column_name_1, column_name_2, ...],
    'data': [
        [cell_value_row_1_col_1, cell_value_row_1_col_2, ...],
        [cell_value_row_2_col_1, cell_value_row_2_col_2, ...],
        ...
    ]
}

解决方案

错误原因是一次性处理全量文本导致内存溢出,CamemBERT处理大批次数据时会占用大量CPU/GPU内存,可通过以下方式解决:

  • 分批处理文本:将数据拆分成小批次逐批处理,避免一次性加载所有数据到内存。调整batch_size参数适配你的内存容量:
import torch
from transformers import AutoTokenizer, CamembertModel
from sklearn.neighbors import KNeighborsClassifier
import numpy as np

tokenizer = AutoTokenizer.from_pretrained("camembert-base")
model = CamembertModel.from_pretrained("camembert-base")

data = df.to_dict(orient='split')
data = dict(zip(data['index'], data['data']))

input_texts = [str(text) for text in data.values()]
input_labels = list(data.keys())

# 设置批次大小,可根据内存情况调整(如32、64)
batch_size = 32
all_features = []

model.eval()
with torch.no_grad():
    for i in range(0, len(input_texts), batch_size):
        # 取当前批次的文本
        batch_texts = input_texts[i:i+batch_size]
        # 分词处理
        inputs = tokenizer(batch_texts, return_tensors="pt", padding=True, truncation=True)
        # 模型推理
        outputs = model(**inputs)
        # 提取<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]>token的特征向量
        batch_features = outputs.last_hidden_state[:, 0, :].numpy()
        all_features.append(batch_features)

# 合并所有批次的特征
input_features = np.concatenate(all_features, axis=0)

# 训练KNN分类器
neigh = KNeighborsClassifier(n_neighbors=3)
neigh.fit(input_features, input_labels)
  • 利用GPU加速:如果有可用GPU,将模型和数据转移到GPU上,GPU内存更适合处理大张量:
# 初始化后添加设备配置
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)

# 批次处理时将输入张量移到对应设备
inputs = {k: v.to(device) for k, v in inputs.items()}
  • 限制文本序列长度:通过max_length参数缩短token序列长度,减少每个样本的内存占用:
inputs = tokenizer(batch_texts, return_tensors="pt", padding=True, truncation=True, max_length=128)
  • 手动清理内存:每个批次处理完成后,删除无用张量并清理缓存:
# 处理完一个批次后执行
del inputs, outputs, batch_features
if torch.cuda.is_available():
    torch.cuda.empty_cache()

内容的提问来源于stack exchange,提问作者Wajih101

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 10:03:19