如何在Python中同时移除英法双语停用词优化聚类效果
Got it, let's tackle this problem head-on—those useless clusters full of French stopwords are definitely messing with your results, and merging the stopword lists is simpler than you might think. Here's how to get it working properly:
Step 1: Combine English and French Stopwords Correctly
The key here is that nltk.corpus.stopwords.words() returns a list for each language. You can just add these lists together, and converting to a set is even better for faster lookups (set membership checks are way quicker than list checks, especially with larger datasets):
import nltk nltk.download('stopwords') # Ensure you've downloaded the stopword datasets first # Load individual language stopword lists english_stop = nltk.corpus.stopwords.words('english') french_stop = nltk.corpus.stopwords.words('french') # Merge into a single set (automatically handles rare cross-language duplicates) combined_stopwords = set(english_stop + french_stop)
If you prefer to keep it as a list for specific workflows, english_stop + french_stop works too—but sets are the better choice for performance when filtering tokens.
Step 2: Filter Tokens Against the Combined List
The reason your previous dual-variable attempt failed is likely due to incorrect filter logic. Instead of checking against two separate lists, use the combined set in your preprocessing step. Here's a straightforward preprocessing function that strips out all bilingual stopwords:
def clean_text(text): # Lowercase first (critical for matching stopwords consistently) lower_text = text.lower() # Split into tokens (adjust this if using a more robust tokenizer like spaCy) tokens = lower_text.split() # Filter out stopwords and non-alphabetic tokens (removes punctuation/numbers) filtered_tokens = [token for token in tokens if token.isalpha() and token not in combined_stopwords] return filtered_tokens
Why This Works
When you run your text through this function, both English stopwords (like "the", "and") and French stopwords (like "la", "et", "dans") will get stripped out entirely. No more clusters dominated by those meaningless terms.
Quick Validation Test
To confirm it's working, test with the exact stopwords from your problematic Cluster 5:
test_cluster_tokens = ["la", "et", "dans", "les", "des", "est", "du", "le", "une", "en", "paris", "london"] cleaned = clean_text(" ".join(test_cluster_tokens)) print(cleaned) # Output should be ['paris', 'london']
Bonus Tip: Robust Tokenization (Optional)
If your text has complex structure (like accented French characters, contractions, etc.), consider using a bilingual tokenizer like spaCy instead of simple split(). You can load both English and French models, merge their stopword sets, and get more accurate tokenization:
import spacy # Install models first: pip install spacy && python -m spacy download en_core_web_sm fr_core_news_sm nlp_en = spacy.load('en_core_web_sm') nlp_fr = spacy.load('fr_core_news_sm') # Merge spaCy's stopword sets spacy_combined_stop = nlp_en.Defaults.stop_words.union(nlp_fr.Defaults.stop_words) # SpaCy-powered cleaning function def spacy_clean(text): # Simplified language detection (adjust based on your dataset needs) doc = nlp_en(text) if any(word in english_stop for word in text.lower().split()) else nlp_fr(text) return [token.text.lower() for token in doc if token.text.lower() not in spacy_combined_stop and token.is_alpha]
This handles edge cases better than basic splitting, but the nltk method is fully sufficient for fixing your immediate clustering issue.
内容的提问来源于stack exchange,提问作者OnThaRise

