You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中同时移除英法双语停用词优化聚类效果

Fixing Bilingual Stopword Removal for Your French-English Text Clustering Project

Got it, let's tackle this problem head-on—those useless clusters full of French stopwords are definitely messing with your results, and merging the stopword lists is simpler than you might think. Here's how to get it working properly:

Step 1: Combine English and French Stopwords Correctly

The key here is that nltk.corpus.stopwords.words() returns a list for each language. You can just add these lists together, and converting to a set is even better for faster lookups (set membership checks are way quicker than list checks, especially with larger datasets):

import nltk
nltk.download('stopwords')  # Ensure you've downloaded the stopword datasets first

# Load individual language stopword lists
english_stop = nltk.corpus.stopwords.words('english')
french_stop = nltk.corpus.stopwords.words('french')

# Merge into a single set (automatically handles rare cross-language duplicates)
combined_stopwords = set(english_stop + french_stop)

If you prefer to keep it as a list for specific workflows, english_stop + french_stop works too—but sets are the better choice for performance when filtering tokens.

Step 2: Filter Tokens Against the Combined List

The reason your previous dual-variable attempt failed is likely due to incorrect filter logic. Instead of checking against two separate lists, use the combined set in your preprocessing step. Here's a straightforward preprocessing function that strips out all bilingual stopwords:

def clean_text(text):
    # Lowercase first (critical for matching stopwords consistently)
    lower_text = text.lower()
    # Split into tokens (adjust this if using a more robust tokenizer like spaCy)
    tokens = lower_text.split()
    # Filter out stopwords and non-alphabetic tokens (removes punctuation/numbers)
    filtered_tokens = [token for token in tokens if token.isalpha() and token not in combined_stopwords]
    return filtered_tokens

Why This Works

When you run your text through this function, both English stopwords (like "the", "and") and French stopwords (like "la", "et", "dans") will get stripped out entirely. No more clusters dominated by those meaningless terms.

Quick Validation Test

To confirm it's working, test with the exact stopwords from your problematic Cluster 5:

test_cluster_tokens = ["la", "et", "dans", "les", "des", "est", "du", "le", "une", "en", "paris", "london"]
cleaned = clean_text(" ".join(test_cluster_tokens))
print(cleaned)  # Output should be ['paris', 'london']

Bonus Tip: Robust Tokenization (Optional)

If your text has complex structure (like accented French characters, contractions, etc.), consider using a bilingual tokenizer like spaCy instead of simple split(). You can load both English and French models, merge their stopword sets, and get more accurate tokenization:

import spacy

# Install models first: pip install spacy && python -m spacy download en_core_web_sm fr_core_news_sm
nlp_en = spacy.load('en_core_web_sm')
nlp_fr = spacy.load('fr_core_news_sm')

# Merge spaCy's stopword sets
spacy_combined_stop = nlp_en.Defaults.stop_words.union(nlp_fr.Defaults.stop_words)

# SpaCy-powered cleaning function
def spacy_clean(text):
    # Simplified language detection (adjust based on your dataset needs)
    doc = nlp_en(text) if any(word in english_stop for word in text.lower().split()) else nlp_fr(text)
    return [token.text.lower() for token in doc if token.text.lower() not in spacy_combined_stop and token.is_alpha]

This handles edge cases better than basic splitting, but the nltk method is fully sufficient for fixing your immediate clustering issue.


内容的提问来源于stack exchange,提问作者OnThaRise

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:29:14