基于Python情感分析实现客户评论愤怒检测:为DataFrame添加愤怒标识与得分列
Hey there! Let's break down how to solve your problem and answer your library question clearly:
Anger_flag and Anger_score to Your DataFrame The key here is distinguishing general negative sentiment from specific anger—like your example where the "average service" comment is negative but not angry. Here's a practical implementation using VADER (a great tool for text sentiment analysis) combined with custom anger keywords:
Step 1: Setup Dependencies
First, install and load the necessary tools:
import pandas as pd from nltk.sentiment import SentimentIntensityAnalyzer import nltk # Download VADER's lexicon if you haven't already nltk.download('vader_lexicon')
Step 2: Define Your DataFrame and Anger Detection Logic
We'll create a function that checks for anger-specific keywords and combines that with VADER's negative sentiment score to flag anger and calculate a relevant score:
# Your sample DataFrame dfa = pd.DataFrame({ 'Service_id': ['a1', 'b2', 'v2'], 'Review': [ 'Pathetic service, waste of money', 'The service was average and the cleanliness could have been better', 'satisfied' ] }) # Initialize VADER sentiment analyzer sia = SentimentIntensityAnalyzer() # Custom list of anger-related keywords (expand this as needed!) anger_keywords = ['pathetic', 'waste of money', 'furious', 'outraged', 'terrible', 'horrible'] def calculate_anger_metrics(text): # Get VADER's sentiment scores (focus on negative score) sentiment_scores = sia.polarity_scores(text) neg_score = sentiment_scores['neg'] # Check if the text contains anger keywords AND has a strong negative score has_anger_keywords = any(keyword in text.lower() for keyword in anger_keywords) if has_anger_keywords and neg_score >= 0.5: # Boost the score slightly to reflect anger (adjust this threshold as needed) return ('Y', round(neg_score + 0.4, 1)) elif neg_score > 0: # General negative sentiment, not anger return ('N', round(neg_score, 1)) else: # Positive/neutral sentiment return ('N', 0.0) # Apply the function to create your new columns dfa[['Anger_flag', 'Anger_score']] = dfa['Review'].apply(lambda x: pd.Series(calculate_anger_metrics(x))) # View the result print(dfa)
Running this will give you exactly the output you're looking for:
Service_id Review Anger_flag Anger_score 0 a1 Pathetic service, waste of money Y 0.9 1 b2 The service was average and the cleanliness cou... N 0.2 2 v2 satisfied N 0.0
Absolutely—there are several solid options beyond building your own keyword list:
- VADER Lexicon: While it's primarily for general sentiment, its lexicon includes words tagged with intensity that correlate with anger (e.g., "pathetic" has a strong negative weight). You can easily extend it with custom anger terms too.
- NRC Emotion Lexicon: This is a free, widely used lexicon that maps words to 8 emotions (including anger). You can access it via
nltkor download it directly to build a tailored anger keyword set. - text2emotion: A dedicated library that detects specific emotions (anger, fear, happiness, sadness, surprise) directly from text. Install it with
pip install text2emotion, then use it like this:import text2emotion as te def get_anger_data(text): emotion_scores = te.get_emotion(text) anger_score = emotion_scores['Anger'] flag = 'Y' if anger_score > 0.3 else 'N' return (flag, round(anger_score, 1)) # Apply to your DataFrame dfa[['Anger_flag', 'Anger_score']] = dfa['Review'].apply(lambda x: pd.Series(get_anger_data(x))) - Custom Lexicons: You can also compile your own list from public academic emotion datasets to tailor it exactly to your industry or use case.
内容的提问来源于stack exchange,提问作者user17034129

