You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用CountVectorizer处理列表列生成词频矩阵时运行崩溃求助

解决大列表型DataFrame生成稀疏词频矩阵的问题

Hey there! Let's work through this problem together—you're hitting a classic memory and speed bottleneck when dealing with large datasets containing list-based category columns. Let's break down why your current methods are failing, then walk through efficient, memory-friendly solutions.

Why Your Current Methods Aren't Working

Let's start by diagnosing the issues with your existing approaches:

  • CountVectorizer崩溃: The biggest problem here is that you're converting the sparse output of fit_transform() to a dense array with toarray(). With 400k rows and 2000+ terms, that's 800 million values stored in memory—way too much for most systems. Also, CountVectorizer is designed for string inputs (like raw text), not lists of terms, which adds unnecessary overhead.
  • 循环方法过慢且分块失效: Nested loops in pandas are inherently slow because they don't leverage vectorized operations. Plus, the deprecated .ix indexer is likely causing issues with non-sequential indexes in your split chunks, leading to failed assignments.

Solution 1: Explode + Crosstab (Pandas-Native, Easy to Implement)

This method uses pandas' built-in functions to avoid loops and keep memory usage low by leveraging sparse data types.

# 1. 先将列表列拆分为每行一个元素(explode会保留原ID关联)
df_exploded = df.explode('Type 1 Category')

# 2. 生成词频交叉表:行=ID,列=术语,值=该ID下术语出现次数
freq_matrix = pd.crosstab(df_exploded['ID'], df_exploded['Type 1 Category'])

# 3. 转换为稀疏格式,大幅节省内存(fill_value=0表示未出现的术语用0填充但不实际存储)
freq_matrix_sparse = freq_matrix.astype(pd.SparseDtype(int, fill_value=0))

# 4. 与原DataFrame合并(确保保留Amount、Age、Fraud等其他列)
final_df = df[['ID', 'Amount', 'Age', 'Fraud']].merge(freq_matrix_sparse, on='ID', how='left')

Solution 2: Scipy Sparse Matrix + Counter (More Flexible for Custom Logic)

If you need more control over term counting (like handling multiple category columns at once), this approach uses scipy's sparse matrices to build the frequency table directly without dense intermediates.

from collections import Counter
import scipy.sparse as sp

# 1. 获取所有唯一术语并建立索引映射
all_terms = df['Type 1 Category'].explode().unique()
term_to_idx = {term: idx for idx, term in enumerate(all_terms)}

# 2. 初始化稀疏矩阵(dok_matrix适合逐元素赋值)
n_rows = len(df)
n_cols = len(all_terms)
sparse_mat = sp.dok_matrix((n_rows, n_cols), dtype=int)

# 3. 逐行统计词频并填充稀疏矩阵
for row_idx, terms_list in enumerate(df['Type 1 Category']):
    term_counts = Counter(terms_list)
    for term, count in term_counts.items():
        sparse_mat[row_idx, term_to_idx[term]] = count

# 4. 转换为Pandas稀疏DataFrame并与原数据合并
freq_df = pd.DataFrame.sparse.from_spmatrix(sparse_mat, index=df.index, columns=all_terms)
final_df = df.drop('Type 1 Category', axis=1).join(freq_df)

Bonus: Fixing Your Original CountVectorizer Approach

If you still want to use CountVectorizer, you can adapt it to handle lists and avoid dense arrays:

from sklearn.feature_extraction.text import CountVectorizer

# 将列表转换为空格分隔的字符串(CountVectorizer的预期输入格式)
df['Type 1 Str'] = df['Type 1 Category'].apply(lambda x: ' '.join(x))

# 初始化CountVectorizer,直接保留稀疏矩阵
countvec = CountVectorizer(vocabulary=all_terms)
sparse_output = countvec.fit_transform(df['Type 1 Str'])

# 转换为稀疏DataFrame
freq_df = pd.DataFrame.sparse.from_spmatrix(sparse_output, index=df.index, columns=countvec.get_feature_names_out())
final_df = df.drop(['Type 1 Category', 'Type 1 Str'], axis=1).join(freq_df)

Additional Tips for Large Datasets

  • 处理多类别列: 如果需要同时处理Type 1/2/3 Categories,可以先把三个列的列表合并成一个列(df['All Types'] = df['Type 1 Category'] + df['Type 2 Category'] + df['Type 3 Category'])再用上述方法。
  • 超大数据集: 如果400k行还是超出内存,试试用Dask——它支持分块处理,语法和Pandas几乎一致,能自动处理内存压力。

内容的提问来源于stack exchange,提问作者wenger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:51:38