使用CamemBERT提取特征后用KNN分类时出现内存不足RuntimeError
CamemBERT特征提取+KNN分类内存不足问题解决
问题场景
我尝试用CamemBERT模型提取文本特征,再用KNN分类器对特征向量分类,编写的代码如下:
import torch from transformers import AutoTokenizer, CamembertModel from sklearn.neighbors import KNeighborsClassifier tokenizer = AutoTokenizer.from_pretrained("camembert-base") model = CamembertModel.from_pretrained("camembert-base") data = df.to_dict(orient='split') data = dict(zip(data['index'], data['data'])) # Collect all the input texts into a list of strings input_texts = [str(text) for text in data.values()] # Tokenize all the input texts together inputs = tokenizer(input_texts, return_tensors="pt", padding=True, truncation=True) # Get the model outputs for all the input texts with torch.no_grad(): outputs = model(**inputs) # Extract the last hidden states and convert them to a numpy array last_hidden_states = outputs.last_hidden_state input_features = last_hidden_states[:, 0, :].numpy() # Extract the labels from the data dictionary input_labels = list(data.keys()) neigh = KNeighborsClassifier(n_neighbors=3) neigh.fit(input_features, input_labels)
报错信息
运行代码后出现内存不足错误:
RuntimeError: [enforce fail at ..\c10\core\impl\alloc_cpu.cpp:72] data. DefaultCPUAllocator: not enough memory: you tried to allocate 19209424896 bytes.
数据格式
使用的数据字典格式如下:
{ 'index': [row_index_1, row_index_2, ...], 'columns': [column_name_1, column_name_2, ...], 'data': [ [cell_value_row_1_col_1, cell_value_row_1_col_2, ...], [cell_value_row_2_col_1, cell_value_row_2_col_2, ...], ... ] }
解决方案
错误原因是一次性处理全量文本导致内存溢出,CamemBERT处理大批次数据时会占用大量CPU/GPU内存,可通过以下方式解决:
- 分批处理文本:将数据拆分成小批次逐批处理,避免一次性加载所有数据到内存。调整
batch_size参数适配你的内存容量:
import torch from transformers import AutoTokenizer, CamembertModel from sklearn.neighbors import KNeighborsClassifier import numpy as np tokenizer = AutoTokenizer.from_pretrained("camembert-base") model = CamembertModel.from_pretrained("camembert-base") data = df.to_dict(orient='split') data = dict(zip(data['index'], data['data'])) input_texts = [str(text) for text in data.values()] input_labels = list(data.keys()) # 设置批次大小,可根据内存情况调整(如32、64) batch_size = 32 all_features = [] model.eval() with torch.no_grad(): for i in range(0, len(input_texts), batch_size): # 取当前批次的文本 batch_texts = input_texts[i:i+batch_size] # 分词处理 inputs = tokenizer(batch_texts, return_tensors="pt", padding=True, truncation=True) # 模型推理 outputs = model(**inputs) # 提取<[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]>token的特征向量 batch_features = outputs.last_hidden_state[:, 0, :].numpy() all_features.append(batch_features) # 合并所有批次的特征 input_features = np.concatenate(all_features, axis=0) # 训练KNN分类器 neigh = KNeighborsClassifier(n_neighbors=3) neigh.fit(input_features, input_labels)
- 利用GPU加速:如果有可用GPU,将模型和数据转移到GPU上,GPU内存更适合处理大张量:
# 初始化后添加设备配置 device = torch.device("cuda" if torch.cuda.is_available() else "cpu") model.to(device) # 批次处理时将输入张量移到对应设备 inputs = {k: v.to(device) for k, v in inputs.items()}
- 限制文本序列长度:通过
max_length参数缩短token序列长度,减少每个样本的内存占用:
inputs = tokenizer(batch_texts, return_tensors="pt", padding=True, truncation=True, max_length=128)
- 手动清理内存:每个批次处理完成后,删除无用张量并清理缓存:
# 处理完一个批次后执行 del inputs, outputs, batch_features if torch.cuda.is_available(): torch.cuda.empty_cache()
内容的提问来源于stack exchange,提问作者Wajih101
相关产品推荐
相关产品推荐

