如何在Python中从CSR矩阵生成向量?大样本BoW训练NN内存优化
处理大尺寸CSR词袋矩阵的神经网络训练方案
看起来你在处理一个典型的大规模文本分类场景——100k样本+100k词汇量的CSR矩阵确实容易把内存撑爆,直接转稠密矩阵肯定行不通。下面给你几个实用的解决方案,结合你的代码场景来调整:
1. 用批量生成器分批加载数据
核心思路是不一次性载入所有数据,而是每次只加载一小批样本转成可用的格式(稠密或稀疏)喂给模型。你可以自定义一个生成器函数,或者用Keras的Sequence类(更安全,支持多线程)。
自定义生成器示例
import numpy as np def csr_batch_generator(csr_matrix, labels, batch_size, shuffle=True): num_samples = csr_matrix.shape[0] indices = np.arange(num_samples) while True: if shuffle: # 每轮打乱样本顺序 np.random.shuffle(indices) for start_idx in range(0, num_samples, batch_size): end_idx = min(start_idx + batch_size, num_samples) batch_indices = indices[start_idx:end_idx] # 提取批次CSR矩阵,按需转稠密/保持稀疏 batch_x = csr_matrix[batch_indices].toarray() # 如果模型支持稀疏张量可以跳过toarray batch_y = labels[batch_indices] yield batch_x, batch_y
训练时使用生成器
# 假设你的模型已经定义好 batch_size = 32 steps_per_epoch = len(train_X) // batch_size model.fit( csr_batch_generator(train_X, train_y, batch_size), steps_per_epoch=steps_per_epoch, epochs=10, validation_data=csr_batch_generator(test_X, test_y, batch_size, shuffle=False), validation_steps=len(test_X) // batch_size )
2. 直接利用框架的稀疏张量支持
像TensorFlow/Keras本身就支持稀疏张量输入,不用把CSR转成稠密矩阵,能省大量内存。
CSR转TensorFlow稀疏张量
import tensorflow as tf def csr_to_tf_sparse(csr_mat): # 获取非零元素的索引、值和矩阵形状 indices = np.column_stack(csr_mat.nonzero()) values = csr_mat.data dense_shape = csr_mat.shape # 转成tf.SparseTensor,注意索引要转成int64 return tf.SparseTensor( indices=tf.cast(indices, tf.int64), values=tf.cast(values, tf.float32), dense_shape=dense_shape )
你可以把这个转换逻辑整合到生成器里,直接返回稀疏张量给模型,模型的输入层要对应设置sparse=True:
input_layer = tf.keras.layers.Input(shape=(100000,), sparse=True) # 后续层正常定义...
3. 先做特征降维(可选但推荐)
100k的词汇量其实有点冗余,你可以先通过特征选择或降维把维度降下来,从根源上减少内存压力:
- 特征选择:用
SelectKBest保留TF-IDF权重最高的Top-N词 - 降维:用
TruncatedSVD对稀疏矩阵做降维(比PCA更适合稀疏数据)
TruncatedSVD降维示例
from sklearn.decomposition import TruncatedSVD # 把100k维降到2000维(可根据任务调整) svd = TruncatedSVD(n_components=2000, random_state=42) train_X_reduced = svd.fit_transform(train_X) test_X_reduced = svd.transform(test_X) # 现在train_X_reduced是稠密矩阵,但维度只有2000,内存压力小很多
顺便纠正你代码里的一个小问题
你当前的代码在每个fold里都对训练数据做fit_transform,但测试集应该用同一个CountVectorizer做transform,而不是重新fit(否则会把测试集的词汇加入词典,导致数据泄露)。正确的写法应该是:
for trainIndices, testIndices in indices.split(X): print("\tAnalyzing fold {}...".format(fold)) print("\t\tIsolating training data...") # 只在训练文本上fit CountVectorizer train_texts = X[trainIndices].apply(lambda x: ' '.join(x)) countVectorizer.fit(train_texts) train_X = countVectorizer.transform(train_texts) train_y = y[trainIndices] print("\t\tIsolating test data...") # 测试文本用已fit的CountVectorizer做transform test_texts = X[testIndices].apply(lambda x: ' '.join(x)) test_X = countVectorizer.transform(test_texts) test_y = y[testIndices] # 后续训练逻辑...
内容的提问来源于stack exchange,提问作者matthp9
相关产品推荐
相关产品推荐

