You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中从CSR矩阵生成向量?大样本BoW训练NN内存优化

处理大尺寸CSR词袋矩阵的神经网络训练方案

看起来你在处理一个典型的大规模文本分类场景——100k样本+100k词汇量的CSR矩阵确实容易把内存撑爆,直接转稠密矩阵肯定行不通。下面给你几个实用的解决方案,结合你的代码场景来调整:

1. 用批量生成器分批加载数据

核心思路是不一次性载入所有数据,而是每次只加载一小批样本转成可用的格式(稠密或稀疏)喂给模型。你可以自定义一个生成器函数,或者用Keras的Sequence类(更安全,支持多线程)。

自定义生成器示例

import numpy as np

def csr_batch_generator(csr_matrix, labels, batch_size, shuffle=True):
    num_samples = csr_matrix.shape[0]
    indices = np.arange(num_samples)
    
    while True:
        if shuffle:
            # 每轮打乱样本顺序
            np.random.shuffle(indices)
        
        for start_idx in range(0, num_samples, batch_size):
            end_idx = min(start_idx + batch_size, num_samples)
            batch_indices = indices[start_idx:end_idx]
            
            # 提取批次CSR矩阵,按需转稠密/保持稀疏
            batch_x = csr_matrix[batch_indices].toarray()  # 如果模型支持稀疏张量可以跳过toarray
            batch_y = labels[batch_indices]
            
            yield batch_x, batch_y

训练时使用生成器

# 假设你的模型已经定义好
batch_size = 32
steps_per_epoch = len(train_X) // batch_size

model.fit(
    csr_batch_generator(train_X, train_y, batch_size),
    steps_per_epoch=steps_per_epoch,
    epochs=10,
    validation_data=csr_batch_generator(test_X, test_y, batch_size, shuffle=False),
    validation_steps=len(test_X) // batch_size
)

2. 直接利用框架的稀疏张量支持

像TensorFlow/Keras本身就支持稀疏张量输入,不用把CSR转成稠密矩阵,能省大量内存。

CSR转TensorFlow稀疏张量

import tensorflow as tf

def csr_to_tf_sparse(csr_mat):
    # 获取非零元素的索引、值和矩阵形状
    indices = np.column_stack(csr_mat.nonzero())
    values = csr_mat.data
    dense_shape = csr_mat.shape
    
    # 转成tf.SparseTensor,注意索引要转成int64
    return tf.SparseTensor(
        indices=tf.cast(indices, tf.int64),
        values=tf.cast(values, tf.float32),
        dense_shape=dense_shape
    )

你可以把这个转换逻辑整合到生成器里,直接返回稀疏张量给模型,模型的输入层要对应设置sparse=True:

input_layer = tf.keras.layers.Input(shape=(100000,), sparse=True)
# 后续层正常定义...

3. 先做特征降维(可选但推荐)

100k的词汇量其实有点冗余,你可以先通过特征选择或降维把维度降下来,从根源上减少内存压力:

  • 特征选择:用SelectKBest保留TF-IDF权重最高的Top-N词
  • 降维:用TruncatedSVD对稀疏矩阵做降维(比PCA更适合稀疏数据)

TruncatedSVD降维示例

from sklearn.decomposition import TruncatedSVD

# 把100k维降到2000维(可根据任务调整)
svd = TruncatedSVD(n_components=2000, random_state=42)
train_X_reduced = svd.fit_transform(train_X)
test_X_reduced = svd.transform(test_X)

# 现在train_X_reduced是稠密矩阵,但维度只有2000,内存压力小很多

顺便纠正你代码里的一个小问题

你当前的代码在每个fold里都对训练数据做fit_transform,但测试集应该用同一个CountVectorizer做transform,而不是重新fit(否则会把测试集的词汇加入词典,导致数据泄露)。正确的写法应该是:

for trainIndices, testIndices in indices.split(X):
    print("\tAnalyzing fold {}...".format(fold))
    print("\t\tIsolating training data...")
    # 只在训练文本上fit CountVectorizer
    train_texts = X[trainIndices].apply(lambda x: ' '.join(x))
    countVectorizer.fit(train_texts)
    train_X = countVectorizer.transform(train_texts)
    train_y = y[trainIndices]
    
    print("\t\tIsolating test data...")
    # 测试文本用已fit的CountVectorizer做transform
    test_texts = X[testIndices].apply(lambda x: ' '.join(x))
    test_X = countVectorizer.transform(test_texts)
    test_y = y[testIndices]
    # 后续训练逻辑...

内容的提问来源于stack exchange,提问作者matthp9

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:44:25