处理314k文本数据时TSNE运行出现MemoryError问题求助
Hey there, let's break down why you're hitting that MemoryError and fix it step by step.
The core issue here is that final_counts.todense() is trying to convert your sparse Bag of Words matrix (from 314k records) into a dense numpy array—and that's way too big for your memory to handle. A sparse matrix only stores non-zero values, but a dense one will allocate space for every single cell, even the zeros. For 314k rows and potentially tens of thousands of features, that's gigabytes (or even terabytes) of data your system can't hold.
Here are concrete fixes you can implement:
1. Pre-reduce dimensionality with PCA before t-SNE
t-SNE works best on lower-dimensional data anyway, so first use PCA to shrink your sparse matrix to a manageable size (like 50 or 100 dimensions) before feeding it to t-SNE. This avoids ever creating that massive dense matrix.
Modify your code like this:
from sklearn.manifold import TSNE from sklearn.decomposition import PCA from sklearn.feature_extraction.text import CountVectorizer import pandas as pd import seaborn as sn import matplotlib.pyplot as plt import numpy as np # Step 1: Generate sparse BoW matrix (keep it sparse!) count_vect = CountVectorizer(max_features=10000) # Limit features to top 10k to reduce size final_counts = count_vect.fit_transform(final['Text'].values) labels = count_vect.get_feature_names() # Step 2: Use PCA to reduce dimensions to 50 (adjust as needed) pca = PCA(n_components=50, random_state=0) pca_data = pca.fit_transform(final_counts.toarray()) # Now this is a small dense matrix # Step 3: Run t-SNE on the reduced PCA data model = TSNE(n_components=2, random_state=0) tsne_data = model.fit_transform(pca_data) # Rest of your plotting code (fix typo: pd.DataFrame not pd.Dataframe) tsne_data = np.vstack((tsne_data.T, labels)).T tsne_df = pd.DataFrame(data=tsne_data, columns=("D_1", "D_2", "label")) sn.FacetGrid(tsne_df, hue="label", height=6).map(plt.scatter, "D_1", "D_2").add_legend() plt.show()
2. Optimize your Bag of Words matrix to reduce its size
Cut down the number of features in your BoW matrix to make even the dense version manageable:
- Use
max_featuresinCountVectorizerto keep only the top N most frequent words (like 10k or 20k as above) - Add
min_dfparameter to ignore words that appear in fewer than X documents (e.g.,min_df=5to skip rare words) - Consider using
TfidfVectorizerinstead ofCountVectorizer—it often produces more meaningful features and can help reduce noise, which might let you use fewer features overall
3. Use a t-SNE implementation that supports sparse matrices
Some newer tools handle sparse inputs directly, which eliminates the need to convert to a dense matrix. For example, MulticoreTSNE supports sparse matrices and runs faster on multiple cores:
from MulticoreTSNE import MulticoreTSNE model = MulticoreTSNE(n_components=2, random_state=0, n_jobs=-1) tsne_data = model.fit_transform(final_counts) # No need to convert to dense!
(Note: You'll need to install this library first with pip install MulticoreTSNE)
Quick typo fix in your original code
You had pd.Dataframe (lowercase 'f')—it should be pd.DataFrame (uppercase 'F'), otherwise you'll hit an error after fixing the MemoryError.
These changes should get you past the memory issue while still producing meaningful t-SNE visualizations. Remember that t-SNE is computationally expensive for large datasets, so reducing the number of features or using PCA first will also speed up the process significantly.
内容的提问来源于stack exchange,提问作者merklexy

