You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

处理314k文本数据时TSNE运行出现MemoryError问题求助

Hey there, let's break down why you're hitting that MemoryError and fix it step by step.

The core issue here is that final_counts.todense() is trying to convert your sparse Bag of Words matrix (from 314k records) into a dense numpy array—and that's way too big for your memory to handle. A sparse matrix only stores non-zero values, but a dense one will allocate space for every single cell, even the zeros. For 314k rows and potentially tens of thousands of features, that's gigabytes (or even terabytes) of data your system can't hold.

Here are concrete fixes you can implement:


1. Pre-reduce dimensionality with PCA before t-SNE

t-SNE works best on lower-dimensional data anyway, so first use PCA to shrink your sparse matrix to a manageable size (like 50 or 100 dimensions) before feeding it to t-SNE. This avoids ever creating that massive dense matrix.

Modify your code like this:

from sklearn.manifold import TSNE
from sklearn.decomposition import PCA
from sklearn.feature_extraction.text import CountVectorizer
import pandas as pd
import seaborn as sn
import matplotlib.pyplot as plt
import numpy as np

# Step 1: Generate sparse BoW matrix (keep it sparse!)
count_vect = CountVectorizer(max_features=10000)  # Limit features to top 10k to reduce size
final_counts = count_vect.fit_transform(final['Text'].values)
labels = count_vect.get_feature_names()

# Step 2: Use PCA to reduce dimensions to 50 (adjust as needed)
pca = PCA(n_components=50, random_state=0)
pca_data = pca.fit_transform(final_counts.toarray())  # Now this is a small dense matrix

# Step 3: Run t-SNE on the reduced PCA data
model = TSNE(n_components=2, random_state=0)
tsne_data = model.fit_transform(pca_data)

# Rest of your plotting code (fix typo: pd.DataFrame not pd.Dataframe)
tsne_data = np.vstack((tsne_data.T, labels)).T
tsne_df = pd.DataFrame(data=tsne_data, columns=("D_1", "D_2", "label"))

sn.FacetGrid(tsne_df, hue="label", height=6).map(plt.scatter, "D_1", "D_2").add_legend()
plt.show()

2. Optimize your Bag of Words matrix to reduce its size

Cut down the number of features in your BoW matrix to make even the dense version manageable:

  • Use max_features in CountVectorizer to keep only the top N most frequent words (like 10k or 20k as above)
  • Add min_df parameter to ignore words that appear in fewer than X documents (e.g., min_df=5 to skip rare words)
  • Consider using TfidfVectorizer instead of CountVectorizer—it often produces more meaningful features and can help reduce noise, which might let you use fewer features overall

3. Use a t-SNE implementation that supports sparse matrices

Some newer tools handle sparse inputs directly, which eliminates the need to convert to a dense matrix. For example, MulticoreTSNE supports sparse matrices and runs faster on multiple cores:

from MulticoreTSNE import MulticoreTSNE

model = MulticoreTSNE(n_components=2, random_state=0, n_jobs=-1)
tsne_data = model.fit_transform(final_counts)  # No need to convert to dense!

(Note: You'll need to install this library first with pip install MulticoreTSNE)

Quick typo fix in your original code

You had pd.Dataframe (lowercase 'f')—it should be pd.DataFrame (uppercase 'F'), otherwise you'll hit an error after fixing the MemoryError.


These changes should get you past the memory issue while still producing meaningful t-SNE visualizations. Remember that t-SNE is computationally expensive for large datasets, so reducing the number of features or using PCA first will also speed up the process significantly.

内容的提问来源于stack exchange,提问作者merklexy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:37:23