You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLP中基于大类词的χ²特征选择技术咨询

Great question! Let’s walk through exactly how to implement chi-squared (χ²) feature selection for your sentiment analysis task, given your preprocessed corpus. χ² works by measuring how much a word’s presence (or absence) deviates from what we’d expect if the word and sentiment class were independent—perfect for filtering out low-impact terms.

Step 1: Define Your Classes and Metrics

First, clarify the basics to avoid confusion:

  • Your positive class = all preprocessed documents labeled positive; negative class = all preprocessed documents labeled negative.
  • For each word in your filtered vocabulary (after stopword removal, stemming, and frequency pruning), we’ll use document frequency (number of documents a word appears in) instead of raw word frequency. This avoids bias from long documents that repeat the same word excessively.
Step 2: Build a Contingency Table for Each Word

The core of χ² calculation is a 2x2 contingency table that captures a word’s presence across sentiment classes. For any word ( w ):

Present in DocumentNot Present in Document
Positive Class( A )( C )
Negative Class( B )( D )

Where:

  • ( A ): Number of positive documents that contain ( w )
  • ( B ): Number of negative documents that contain ( w )
  • ( C ): Total positive documents - ( A ) (positive docs without ( w ))
  • ( D ): Total negative documents - ( B ) (negative docs without ( w ))
  • ( N = A + B + C + D ): Total number of documents in your corpus
Step 3: Calculate the χ² Statistic

Use this formula to compute how strongly ( w ) is associated with sentiment class:
[
\chi^2 = \frac{N \times (AD - BC)^2}{(A+B)(C+D)(A+C)(B+D)}
]

  • If ( w ) is independent of sentiment (no meaningful association), ( \chi^2 ) will be close to 0.
  • If ( w ) is strongly linked to one class (e.g., "fantastic" mostly in positive docs), ( \chi^2 ) will be large.

Pro Tip: If any value in the table is 0 (e.g., a word never appears in negative docs), add 1 to all four values (Laplace smoothing) to avoid division-by-zero errors.

Step 4: Rank and Select Features

Once you’ve calculated ( \chi^2 ) for every word in your vocabulary:

  1. Sort all words in descending order of their ( \chi^2 ) values.
  2. Select the top ( K ) words (e.g., top 1000, top 20% of vocabulary) as your final feature set. These are the words most strongly associated with positive or negative sentiment.

Alternatively, you can filter words based on a p-value threshold (e.g., keep words where ( p < 0.05 )), but ranking by ( \chi^2 ) score is more common for sentiment tasks since it directly prioritizes discriminative terms.

Practical Implementation Example (Python)

You can use scikit-learn to automate this quickly, or manually calculate if you want full control. Here’s the scikit-learn approach, using binary document frequency (since we care about presence/absence, not count):

from sklearn.feature_selection import SelectKBest, chi2
from sklearn.feature_extraction.text import CountVectorizer

# Assume:
# X = list of preprocessed documents (each string is a cleaned, stemmed text)
# y = list of sentiment labels (0 for negative, 1 for positive)

# Create a binary document-word matrix (1 = word present, 0 = not present)
vectorizer = CountVectorizer(binary=True)
X_binary = vectorizer.fit_transform(X)

# Select top 1000 features using chi-squared
selector = SelectKBest(chi2, k=1000)
X_selected = selector.fit_transform(X_binary, y)

# Get the actual words that were selected
selected_words = [vectorizer.get_feature_names_out()[idx] for idx in selector.get_support(indices=True)]

If you want to manually calculate ( \chi^2 ) for a specific word:

def calculate_chi2(A, B, C, D):
    N = A + B + C + D
    numerator = N * (A*D - B*C)**2
    denominator = (A+B) * (C+D) * (A+C) * (B+D)
    # Handle division by zero with smoothing if needed
    if denominator == 0:
        return 0.0
    return numerator / denominator

# Example values for the word "awesome"
A = 60  # positive docs with "awesome"
B = 10  # negative docs with "awesome"
C = 40  # positive docs without "awesome"
D = 90  # negative docs without "awesome"

chi2_score = calculate_chi2(A, B, C, D)
print(f"χ² score for 'awesome': {chi2_score:.2f}")  # Outputs ~54.95 (strong positive association)
Key Notes to Avoid Pitfalls
  • Skip raw word frequency: Raw counts can skew results toward words that repeat a lot in a single long document, rather than words that appear across many documents of one sentiment.
  • Smooth zero values: If a word never appears in one class, smoothing prevents invalid calculations.
  • Validate your feature set: Test how well the selected features perform in your sentiment model (e.g., logistic regression, Naive Bayes) to ensure you’re not overfitting to noise.

内容的提问来源于stack exchange,提问作者Ben

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:45:33