基于Python的Twitter数据情感分析:循环代码逻辑问题咨询
Hey there! Let's get your Twitter sentiment analysis code sorted out—your core idea is solid, but we need to fix a few critical logic gaps that are throwing off your counts. Here's a breakdown of the issues and a fully revised version of your code:
Key Issues in Your Original Code
- Count variables reset every loop: You initialized
negative,positive, andneutralinside thefor line in floop. That means every time the loop processes a new tweet, those counts get set back to 0—so you'll never accumulate totals across all tweets. - Missing classification assignment: The line that gets the actual sentiment result (
blob.classify()) was commented out, soclassificationnever gets a value. No wonder the counting wasn't working! - Incomplete condition logic: You didn't finish the positive/negative counting branches, and redefined your sentiment labels (
pos,neg,neu) inside the loop unnecessarily.
Revised & Fixed Code
from textblob import TextBlob # Import your classifier type (adjust if you're using a different one) from textblob.classifiers import NaiveBayesClassifier # 1. Initialize count variables OUTSIDE the loop to accumulate totals total_negative = 0 total_positive = 0 total_neutral = 0 # 2. Define sentiment labels once (no need to redefine every loop) POS_LABEL = "pos" NEG_LABEL = "neg" NEU_LABEL = "neu" # IMPORTANT: Make sure your classifier is trained before this point! # Example training (replace with your actual training data): # train_data = [("I love this!", "pos"), ("Terrible experience", "neg"), ("Meh, it's okay", "neu")] # classifier = NaiveBayesClassifier(train_data) with open('POStagged_tweets_lines.json', 'r') as f: for line in f: # Clean up the input line (remove extra newlines/whitespace) cleaned_tweet = line.strip() # 3. Create TextBlob instance and get the sentiment classification blob = TextBlob(cleaned_tweet, classifier=classifier) classification = blob.classify() # This is the key line you were missing # 4. Complete counting logic with all sentiment cases if classification == NEU_LABEL: total_neutral += 1 elif classification == POS_LABEL: total_positive += 1 elif classification == NEG_LABEL: total_negative += 1 else: # Catch unexpected labels to avoid silent failures print(f"Warning: Unknown sentiment label detected: {classification}") # 5. Calculate and print final stats (matches your note about 90% neutral) total_tweets = total_negative + total_positive + total_neutral print(f"=== Sentiment Analysis Results ===") print(f"Total tweets analyzed: {total_tweets}") print(f"Neutral: {total_neutral} ({(total_neutral / total_tweets) * 100:.2f}%)") print(f"Positive: {total_positive} ({(total_positive / total_tweets) * 100:.2f}%)") print(f"Negative: {total_negative} ({(total_negative / total_tweets) * 100:.2f}%)")
Extra Notes for Your Use Case
- JSON Data Handling: If your
POStagged_tweets_lines.jsonhas structured JSON objects per line (not just raw text), you'll need to parse each line first withjson.loads()to extract the tweet text. For example:import json # ... inside the loop: tweet_obj = json.loads(line) cleaned_tweet = tweet_obj['text'].strip() # Replace 'text' with your actual field name - Classifier Neutral Support: Since you mentioned 90% of your tags are neutral, double-check that your training data includes neutral examples. If your classifier wasn't trained on neutral sentiment, it won't output
neulabels—this is crucial to getting accurate counts.
内容的提问来源于stack exchange,提问作者bybu
相关产品推荐
相关产品推荐

