基于Python计算两本TXT书籍的TF-IDF词相关性及报错排查
Hey there, let's tackle your TF-IDF problem step by step. I've broken down solutions for your error, performance issues, and alternative approaches below:
First, Squash That TypeError: coercing to Unicode: need string or buffer, float found
This error almost always means you're passing a float value to a function that expects text strings. It's likely happening during text loading/preprocessing—maybe empty lines got parsed as NaN (a float) or a preprocessing step returned a non-string value.
Try this robust text loading/preprocessing template to ensure every line is a valid string:
import string def load_and_clean_text(file_path): with open(file_path, 'r', encoding='utf-8') as f: # Filter out empty lines and strip whitespace lines = [line.strip() for line in f if line.strip()] # Add basic preprocessing: lowercase + remove punctuation cleaned_lines = [ line.lower().translate(str.maketrans('', '', string.punctuation)) for line in lines ] return cleaned_lines # Load both books book1 = load_and_clean_text('book1.txt') book2 = load_and_clean_text('book2.txt') # Combine into a single corpus (required for IDF calculation) corpus = book1 + book2
Speed Up Calculations & Avoid Crashes
6000+ lines isn't huge, but manual TF-IDF implementations (like the one in the blog you referenced) are slow and memory-heavy. Use scikit-learn's optimized TF-IDF tools—they're battle-tested and way more efficient.
Full Working Code Example
from sklearn.feature_extraction.text import TfidfVectorizer # Use the load_and_clean_text function from above to get your corpus corpus = book1 + book2 # Initialize TF-IDF vectorizer (customize parameters as needed) tfidf_vectorizer = TfidfVectorizer( stop_words='english', # Remove common words like "the", "and" # max_features=10000 # Uncomment if you want to limit top vocabulary size ) # Calculate TF-IDF scores tfidf_matrix = tfidf_vectorizer.fit_transform(corpus) # Extract vocabulary and their total TF-IDF scores (summed across all docs) feature_names = tfidf_vectorizer.get_feature_names_out() total_tfidf_scores = tfidf_matrix.sum(axis=0).A1 # Sort words by score (descending) sorted_words = sorted( zip(feature_names, total_tfidf_scores), key=lambda x: x[1], reverse=True ) # Print top 20 highest-scoring words for word, score in sorted_words[:20]: print(f"{word}: {score:.4f}")
Why This Works Better:
- Scikit-learn uses sparse matrices to save memory (no more crashes from bloated data)
- Underlying code is C-optimized—way faster than pure Python implementations
- Built-in tools for stopword removal, tokenization, and more (less code to debug)
A Quick Note on the Java TF-IDF Approach
You mentioned trying to call a Java TF-IDF program from Python, but cut off mid-thought. Honestly, this adds unnecessary complexity—scikit-learn's solution is more than fast enough for your dataset. If you still want to pursue it, you'd use Python's subprocess module to run the Java jar, but it's overkill here.
Extra Tips for Edge Cases
- If working with Chinese text, first segment sentences into words using a library like
jieba, then pass the segmented text toTfidfVectorizer - Double-check your TXT file encoding (e.g.,
gbkinstead ofutf-8) if you see garbled text during loading - If you need per-book TF-IDF scores instead of corpus-wide, split the matrix and calculate scores per book separately
内容的提问来源于stack exchange,提问作者SctALE

