如何在YouTube评论情感分析中移除无情感倾向的词汇?
Hey there, I get your frustration—relying solely on POS tagging isn’t working because some sentiment-critical words are getting lumped in with neutral proper nouns and common nouns. Let’s walk through a few practical fixes tailored to your use case:
1. Custom Stopword List (Quickest Fix)
Since you already know the exact neutral terms you want to exclude (mrbeast, james, lion, tigers, pewdiepie, etc.), creating a custom stopword list is the fastest way to filter them out without accidentally removing sentiment words like clickbait or fight.
Here’s how to implement this in Python:
# Define your custom neutral terms (lowercase to handle case variations) neutral_terms = {"mrbeast", "james", "lion", "tigers", "pewdiepie", "tiger's", "lion's"} # Input text input_text = "mrbeast james lion tigers bad sad clickbait fight nice good" # Split into words and filter out neutral terms filtered_words = [word for word in input_text.split() if word.lower() not in neutral_terms] # Result: ['bad', 'sad', 'clickbait', 'fight', 'nice', 'good'] print(filtered_words)
This approach is reliable because you’re explicitly targeting the terms you know are neutral, so you won’t lose important sentiment vocabulary.
2. Hybrid POS + Sentiment Lexicon Check
If you still want to use POS tagging to catch unexpected neutral nouns, pair it with a sentiment lexicon to avoid removing words that carry emotional weight. Tools like VADER (Valence Aware Dictionary and sEntiment Reasoner) are perfect for this—it’s built for social media text and understands colloquial terms like clickbait.
Example using NLTK’s VADER:
from nltk.sentiment import SentimentIntensityAnalyzer import nltk nltk.download('vader_lexicon') sia = SentimentIntensityAnalyzer() input_words = input_text.split() neutral_terms = {"mrbeast", "james", "lion", "tigers", "pewdiepie", "tiger's", "lion's"} filtered_words = [] for word in input_words: # Skip known neutral terms first if word.lower() in neutral_terms: continue # Check if the word has non-neutral sentiment sentiment_score = sia.polarity_scores(word)['compound'] if sentiment_score != 0: filtered_words.append(word) # Result: ['bad', 'sad', 'clickbait', 'fight', 'nice', 'good'] print(filtered_words)
VADER will correctly identify that clickbait and fight carry negative sentiment, so they’ll be kept even if POS tagging labels them as NN.
3. Improve POS Tagging for Colloquial Text
The default NLTK Average Perceptron Tagger is trained on formal text, which is why it mislabels colloquial terms like clickbait as NN. If you want to fix the tagging itself, you could:
- Train a custom tagger on a dataset of YouTube comments (or similar informal text) to get more accurate POS labels.
- Use spaCy’s pre-trained models, which are better at handling informal language. For example:
import spacy nlp = spacy.load("en_core_web_sm") doc = nlp(input_text) # Check POS tags with spaCy for token in doc: print((token.text, token.pos_))
Even if spaCy still tags clickbait as a noun, combining this with the sentiment lexicon check from solution 2 will ensure you don’t remove it.
The key takeaway here is that relying solely on POS tagging for this task is risky—sentiment can exist in any part of speech. Pairing explicit neutral term filtering with a sentiment-aware check will give you the best results for YouTube comment analysis.
内容的提问来源于stack exchange,提问作者Amazing World

