使用DataFrame.apply()时出现语法错误,求代码修正方案
Fixing Your VADER Sentiment Analysis Syntax Error & Code Issues
Let's break down the problems in your code and fix them step by step.
1. Root Cause of the SyntaxError
Your lambda line has two critical syntax mistakes:
- Extra closing bracket:
scores["compound"]]should bescores["compound"](the extra]is invalid). - Missing
elsein the ternary condition chain: Python requireselsebetween consecutiveifclauses in a ternary expression.
2. Broader Issues in Your Code Structure
Beyond the syntax error, your code has inefficiencies and logical flaws:
- You're overwriting your DataFrame inside the loop with
df = pd.DataFrame(columns=['sentiment']), which erases your original data. - Looping through the
bodycolumn as a list is unnecessary—pandas'applymethod can handle this directly on the column. - Applying the lambda to an empty
sentimentcolumn doesn't serve any purpose; you need to compute sentiment from thebodytext and assign it to the new column.
3. Corrected Code
Here's the cleaned-up, working version (I'll use a named function for readability, though a lambda is also possible):
import pandas as pd # Don't forget to import pandas! import nltk nltk.download('vader_lexicon') nltk.download('punkt') from nltk.sentiment.vader import SentimentIntensityAnalyzer # Initialize VADER once (no need to reinitialize in loops) sid = SentimentIntensityAnalyzer() # Define a helper function to calculate sentiment def determine_sentiment(text): sentiment_scores = sid.polarity_scores(text) compound_score = sentiment_scores['compound'] if compound_score >= 0.05: return 'positive' elif compound_score <= -0.05: return 'negative' else: return 'neutral' # Apply the function to your 'body' column to create the 'sentiment' column df['sentiment'] = df['body'].apply(determine_sentiment)
If You Prefer Using a Lambda
If you want to stick with a lambda (though less readable for complex logic), here's the corrected version:
df['sentiment'] = df['body'].apply( lambda text: 'positive' if sid.polarity_scores(text)['compound'] >= 0.05 else 'negative' if sid.polarity_scores(text)['compound'] <= -0.05 else 'neutral' )
Key Improvements
- No more loops: Using
applydirectly on thebodycolumn is more efficient and idiomatic pandas. - Preserves original data: We modify the existing DataFrame instead of creating a new empty one each time.
- Readable logic: The named function makes it easy to adjust sentiment thresholds or add additional logic later.
内容的提问来源于stack exchange,提问作者ASH
相关产品推荐
相关产品推荐

