NLP中基于大类词的χ²特征选择技术咨询
Great question! Let’s walk through exactly how to implement chi-squared (χ²) feature selection for your sentiment analysis task, given your preprocessed corpus. χ² works by measuring how much a word’s presence (or absence) deviates from what we’d expect if the word and sentiment class were independent—perfect for filtering out low-impact terms.
First, clarify the basics to avoid confusion:
- Your positive class = all preprocessed documents labeled positive; negative class = all preprocessed documents labeled negative.
- For each word in your filtered vocabulary (after stopword removal, stemming, and frequency pruning), we’ll use document frequency (number of documents a word appears in) instead of raw word frequency. This avoids bias from long documents that repeat the same word excessively.
The core of χ² calculation is a 2x2 contingency table that captures a word’s presence across sentiment classes. For any word ( w ):
| Present in Document | Not Present in Document | |
|---|---|---|
| Positive Class | ( A ) | ( C ) |
| Negative Class | ( B ) | ( D ) |
Where:
- ( A ): Number of positive documents that contain ( w )
- ( B ): Number of negative documents that contain ( w )
- ( C ): Total positive documents - ( A ) (positive docs without ( w ))
- ( D ): Total negative documents - ( B ) (negative docs without ( w ))
- ( N = A + B + C + D ): Total number of documents in your corpus
Use this formula to compute how strongly ( w ) is associated with sentiment class:
[
\chi^2 = \frac{N \times (AD - BC)^2}{(A+B)(C+D)(A+C)(B+D)}
]
- If ( w ) is independent of sentiment (no meaningful association), ( \chi^2 ) will be close to 0.
- If ( w ) is strongly linked to one class (e.g., "fantastic" mostly in positive docs), ( \chi^2 ) will be large.
Pro Tip: If any value in the table is 0 (e.g., a word never appears in negative docs), add 1 to all four values (Laplace smoothing) to avoid division-by-zero errors.
Once you’ve calculated ( \chi^2 ) for every word in your vocabulary:
- Sort all words in descending order of their ( \chi^2 ) values.
- Select the top ( K ) words (e.g., top 1000, top 20% of vocabulary) as your final feature set. These are the words most strongly associated with positive or negative sentiment.
Alternatively, you can filter words based on a p-value threshold (e.g., keep words where ( p < 0.05 )), but ranking by ( \chi^2 ) score is more common for sentiment tasks since it directly prioritizes discriminative terms.
You can use scikit-learn to automate this quickly, or manually calculate if you want full control. Here’s the scikit-learn approach, using binary document frequency (since we care about presence/absence, not count):
from sklearn.feature_selection import SelectKBest, chi2 from sklearn.feature_extraction.text import CountVectorizer # Assume: # X = list of preprocessed documents (each string is a cleaned, stemmed text) # y = list of sentiment labels (0 for negative, 1 for positive) # Create a binary document-word matrix (1 = word present, 0 = not present) vectorizer = CountVectorizer(binary=True) X_binary = vectorizer.fit_transform(X) # Select top 1000 features using chi-squared selector = SelectKBest(chi2, k=1000) X_selected = selector.fit_transform(X_binary, y) # Get the actual words that were selected selected_words = [vectorizer.get_feature_names_out()[idx] for idx in selector.get_support(indices=True)]
If you want to manually calculate ( \chi^2 ) for a specific word:
def calculate_chi2(A, B, C, D): N = A + B + C + D numerator = N * (A*D - B*C)**2 denominator = (A+B) * (C+D) * (A+C) * (B+D) # Handle division by zero with smoothing if needed if denominator == 0: return 0.0 return numerator / denominator # Example values for the word "awesome" A = 60 # positive docs with "awesome" B = 10 # negative docs with "awesome" C = 40 # positive docs without "awesome" D = 90 # negative docs without "awesome" chi2_score = calculate_chi2(A, B, C, D) print(f"χ² score for 'awesome': {chi2_score:.2f}") # Outputs ~54.95 (strong positive association)
- Skip raw word frequency: Raw counts can skew results toward words that repeat a lot in a single long document, rather than words that appear across many documents of one sentiment.
- Smooth zero values: If a word never appears in one class, smoothing prevents invalid calculations.
- Validate your feature set: Test how well the selected features perform in your sentiment model (e.g., logistic regression, Naive Bayes) to ensure you’re not overfitting to noise.
内容的提问来源于stack exchange,提问作者Ben

