sklearn/nltk中短语停用词被忽略的技术问题咨询
Got it, let's break this down. The root problem here is that multi-word stopwords (like your example company name, e.g., "Acme Corporation") get split into individual tokens when you run tokenization first. By the time you go to filter stopwords, you're checking single tokens against multi-word phrases—so they never match, and those terms slip through.
Here's how to fix this by reordering your pipeline to handle multi-word stopwords before you split text into tokens, plus a concrete implementation using your existing tools (nltk, sklearn):
Step 1: Adjust Your Pipeline Order
Your original flow is:
Tokenize → Lemmatize → Filter stopwords → Generate n-grams
We need to rework it to:
Remove multi-word stopwords → Tokenize → Lemmatize → Filter single-word stopwords → Generate n-grams
Multi-word stopwords are continuous text sequences, so you have to catch them before splitting the text into individual tokens. Once they're split, there's no way to recognize the original phrase.
Step 2: Implement Multi-Word Stopword Removal
Use regular expressions to target and remove multi-word stopwords while preserving word boundaries (so you don't accidentally replace partial matches). Here's the code:
import re import nltk from nltk.stem import WordNetLemmatizer from sklearn.feature_extraction.text import CountVectorizer from sklearn.feature_extraction import text # Download required NLTK resources (run once) nltk.download('wordnet') nltk.download('omw-1.4') # Define your stopwords: single-word + multi-word (e.g., company names) custom_single_stopwords = ["foo", "bar"] # Example single-word stops custom_multi_stopwords = ["Acme Corporation"] # Your multi-word stopword # Merge sklearn's default English stopwords with your single-word stops stop_words = text.ENGLISH_STOP_WORDS.union(custom_single_stopwords) def remove_multi_word_stopwords(raw_text, multi_stop_list): # For each multi-word phrase, build a regex to match the exact phrase (word boundaries included) for phrase in multi_stop_list: # Escape special characters and add word boundaries to avoid partial matches regex_pattern = r'\b' + re.escape(phrase) + r'\b' # Replace the phrase with empty string, ignore case for flexibility raw_text = re.sub(regex_pattern, '', raw_text, flags=re.IGNORECASE) # Clean up extra spaces left from removal return re.sub(r'\s+', ' ', raw_text).strip()
Step 3: Run the Updated Pipeline
Now apply the adjusted flow to your document:
# Example input document sample_doc = "We held a strategy call with Acme Corporation last week. Acme Corporation's team reviewed project foo and bar." # 1. Remove multi-word stopwords first cleaned_doc = remove_multi_word_stopwords(sample_doc, custom_multi_stopwords) # 2. Tokenize using your custom pattern (2+ alphanumeric chars, word-boundary wrapped) token_pattern = r'\b\w{2,}\b' tokens = re.findall(token_pattern, cleaned_doc.lower()) # 3. Lemmatize tokens with NLTK lemmatizer = WordNetLemmatizer() lemmatized_tokens = [lemmatizer.lemmatize(token) for token in tokens] # 4. Filter out single-word stopwords filtered_tokens = [token for token in lemmatized_tokens if token not in stop_words] # 5. Generate 1-4 gram word frequencies vectorizer = CountVectorizer(ngram_range=(1, 4), tokenizer=lambda x: x) freq_matrix = vectorizer.fit_transform([filtered_tokens]) word_freq = dict(zip(vectorizer.get_feature_names_out(), freq_matrix.toarray()[0])) # Check the result (you won't see "acme" or "corporation" anywhere) print(word_freq)
Why This Works
By removing multi-word stopwords upfront, you eliminate the entire phrase before it gets split into individual tokens. This ensures terms like "Acme" and "Corporation" don't even make it to the tokenization step, so you don't have to worry about them slipping through the single-word stopword filter.
Alternative (Less Efficient) Approach: If you absolutely can't reorder the pipeline, you could generate your n-grams first, then filter out any n-gram that matches your multi-word stopwords. But this wastes resources generating unnecessary n-grams, so the upfront removal is always better.
内容的提问来源于stack exchange,提问作者Brad Solomon

