You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python的邮件文本分析:TF-IDF生成主题关键词DataFrame报错排查

Fixing ValueError: Length mismatch when setting TF-IDF feature names as columns

Got it, let's figure out why you're hitting this error and walk through a corrected approach to build your topic-keyword DataFrame.

What's causing the error?

The error ValueError: Length mismatch: Expected axis has 0 elements, new values have 12730166 elements tells us your df_topic_keywords DataFrame is empty (0 columns/rows) when you try to assign 12 million+ feature names as columns. This usually happens because:

  • You didn't filter out empty documents/topics, leading to groups with no valid text to process with TF-IDF.
  • Your initial grouping/TF-IDF logic failed to populate df_topic_keywords with any data before setting columns.
  • You didn't limit TF-IDF features, leading to an absurdly large number of terms that's impractical and causes the mismatch.

Corrected Step-by-Step Implementation

Let's rebuild the code properly, with data cleaning, controlled TF-IDF, and targeted keyword extraction per topic.

First, import required libraries:

import pandas as pd
import re
from sklearn.feature_extraction.text import TfidfVectorizer
from collections import defaultdict

1. Clean your data

First, remove rows with empty documents or topics to avoid processing invalid groups:

# Drop rows where document or topic is missing
df_clean = df.dropna(subset=["document", "topic"]).copy()

# Optional: Preprocess text to improve keyword quality
def preprocess_text(text):
    # Convert to lowercase
    text = text.lower()
    # Remove punctuation, numbers, and extra whitespace
    text = re.sub(r"[^a-zA-Z\s]", "", text)
    text = re.sub(r"\s+", " ", text).strip()
    return text

df_clean["document"] = df_clean["document"].apply(preprocess_text)

2. Extract Top Keywords per Topic with TF-IDF

We'll group by topic, apply TF-IDF to each group, and extract the highest-scoring terms:

# Configure TF-IDF with sensible limits to avoid excessive features
tfidf = TfidfVectorizer(
    stop_words="english",
    ngram_range=(1, 2),  # Include single words and word pairs
    max_features=10000,  # Limit to top 10k most frequent terms
    min_df=2  # Ignore terms that appear in fewer than 2 documents
)

# Store topic-to-keywords mapping
topic_keywords = defaultdict(list)

# Process each topic group
for topic, group in df_clean.groupby("topic"):
    corpus = group["document"].tolist()
    # Skip groups with no valid text after preprocessing
    if not any(corpus):
        continue
    
    # Fit TF-IDF to the topic's corpus
    tfidf_matrix = tfidf.fit_transform(corpus)
    # Get feature names (use get_feature_names_out() for sklearn >= 1.0)
    feature_names = tfidf.get_feature_names_out()
    # Calculate average TF-IDF score for each term across the topic's documents
    avg_tfidf_scores = tfidf_matrix.mean(axis=0).tolist()[0]
    
    # Sort terms by score and pick top N (e.g., top 10 keywords)
    sorted_terms = sorted(zip(feature_names, avg_tfidf_scores), key=lambda x: x[1], reverse=True)[:10]
    # Extract just the keyword strings
    topic_keywords[topic] = [term[0] for term in sorted_terms]

3. Build the Final Topic-Keyword DataFrame

Now convert our mapping into a structured DataFrame with topics as the index:

# Convert the keyword dictionary to a DataFrame
df_topic_keywords = pd.DataFrame.from_dict(topic_keywords, orient="index")

# Rename columns to be descriptive (e.g., keyword_1, keyword_2)
df_topic_keywords.columns = [f"keyword_{i+1}" for i in range(df_topic_keywords.shape[1])]

Key Fixes & Improvements

  • Data Cleaning: We removed invalid rows to ensure every topic group has text to process.
  • Controlled TF-IDF: max_features and min_df prevent the creation of millions of irrelevant terms, keeping processing fast and manageable.
  • Per-Topic TF-IDF: We fit TF-IDF individually per topic, ensuring keywords are relevant to that specific topic (instead of global terms).
  • Avoid Empty DataFrames: We only populate the keyword dictionary with valid topics, so the final DataFrame will never be empty when setting columns.

内容的提问来源于stack exchange,提问作者abombz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:30:21