基于Python的邮件文本分析:TF-IDF生成主题关键词DataFrame报错排查
Got it, let's figure out why you're hitting this error and walk through a corrected approach to build your topic-keyword DataFrame.
What's causing the error?
The error ValueError: Length mismatch: Expected axis has 0 elements, new values have 12730166 elements tells us your df_topic_keywords DataFrame is empty (0 columns/rows) when you try to assign 12 million+ feature names as columns. This usually happens because:
- You didn't filter out empty documents/topics, leading to groups with no valid text to process with TF-IDF.
- Your initial grouping/TF-IDF logic failed to populate
df_topic_keywordswith any data before setting columns. - You didn't limit TF-IDF features, leading to an absurdly large number of terms that's impractical and causes the mismatch.
Corrected Step-by-Step Implementation
Let's rebuild the code properly, with data cleaning, controlled TF-IDF, and targeted keyword extraction per topic.
First, import required libraries:
import pandas as pd import re from sklearn.feature_extraction.text import TfidfVectorizer from collections import defaultdict
1. Clean your data
First, remove rows with empty documents or topics to avoid processing invalid groups:
# Drop rows where document or topic is missing df_clean = df.dropna(subset=["document", "topic"]).copy() # Optional: Preprocess text to improve keyword quality def preprocess_text(text): # Convert to lowercase text = text.lower() # Remove punctuation, numbers, and extra whitespace text = re.sub(r"[^a-zA-Z\s]", "", text) text = re.sub(r"\s+", " ", text).strip() return text df_clean["document"] = df_clean["document"].apply(preprocess_text)
2. Extract Top Keywords per Topic with TF-IDF
We'll group by topic, apply TF-IDF to each group, and extract the highest-scoring terms:
# Configure TF-IDF with sensible limits to avoid excessive features tfidf = TfidfVectorizer( stop_words="english", ngram_range=(1, 2), # Include single words and word pairs max_features=10000, # Limit to top 10k most frequent terms min_df=2 # Ignore terms that appear in fewer than 2 documents ) # Store topic-to-keywords mapping topic_keywords = defaultdict(list) # Process each topic group for topic, group in df_clean.groupby("topic"): corpus = group["document"].tolist() # Skip groups with no valid text after preprocessing if not any(corpus): continue # Fit TF-IDF to the topic's corpus tfidf_matrix = tfidf.fit_transform(corpus) # Get feature names (use get_feature_names_out() for sklearn >= 1.0) feature_names = tfidf.get_feature_names_out() # Calculate average TF-IDF score for each term across the topic's documents avg_tfidf_scores = tfidf_matrix.mean(axis=0).tolist()[0] # Sort terms by score and pick top N (e.g., top 10 keywords) sorted_terms = sorted(zip(feature_names, avg_tfidf_scores), key=lambda x: x[1], reverse=True)[:10] # Extract just the keyword strings topic_keywords[topic] = [term[0] for term in sorted_terms]
3. Build the Final Topic-Keyword DataFrame
Now convert our mapping into a structured DataFrame with topics as the index:
# Convert the keyword dictionary to a DataFrame df_topic_keywords = pd.DataFrame.from_dict(topic_keywords, orient="index") # Rename columns to be descriptive (e.g., keyword_1, keyword_2) df_topic_keywords.columns = [f"keyword_{i+1}" for i in range(df_topic_keywords.shape[1])]
Key Fixes & Improvements
- Data Cleaning: We removed invalid rows to ensure every topic group has text to process.
- Controlled TF-IDF:
max_featuresandmin_dfprevent the creation of millions of irrelevant terms, keeping processing fast and manageable. - Per-Topic TF-IDF: We fit TF-IDF individually per topic, ensuring keywords are relevant to that specific topic (instead of global terms).
- Avoid Empty DataFrames: We only populate the keyword dictionary with valid topics, so the final DataFrame will never be empty when setting columns.
内容的提问来源于stack exchange,提问作者abombz

