You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python的Twitter数据情感分析:循环代码逻辑问题咨询

Hey there! Let's get your Twitter sentiment analysis code sorted out—your core idea is solid, but we need to fix a few critical logic gaps that are throwing off your counts. Here's a breakdown of the issues and a fully revised version of your code:

Key Issues in Your Original Code

  • Count variables reset every loop: You initialized negative, positive, and neutral inside the for line in f loop. That means every time the loop processes a new tweet, those counts get set back to 0—so you'll never accumulate totals across all tweets.
  • Missing classification assignment: The line that gets the actual sentiment result (blob.classify()) was commented out, so classification never gets a value. No wonder the counting wasn't working!
  • Incomplete condition logic: You didn't finish the positive/negative counting branches, and redefined your sentiment labels (pos, neg, neu) inside the loop unnecessarily.

Revised & Fixed Code

from textblob import TextBlob
# Import your classifier type (adjust if you're using a different one)
from textblob.classifiers import NaiveBayesClassifier

# 1. Initialize count variables OUTSIDE the loop to accumulate totals
total_negative = 0
total_positive = 0
total_neutral = 0

# 2. Define sentiment labels once (no need to redefine every loop)
POS_LABEL = "pos"
NEG_LABEL = "neg"
NEU_LABEL = "neu"

# IMPORTANT: Make sure your classifier is trained before this point!
# Example training (replace with your actual training data):
# train_data = [("I love this!", "pos"), ("Terrible experience", "neg"), ("Meh, it's okay", "neu")]
# classifier = NaiveBayesClassifier(train_data)

with open('POStagged_tweets_lines.json', 'r') as f:
    for line in f:
        # Clean up the input line (remove extra newlines/whitespace)
        cleaned_tweet = line.strip()
        
        # 3. Create TextBlob instance and get the sentiment classification
        blob = TextBlob(cleaned_tweet, classifier=classifier)
        classification = blob.classify()  # This is the key line you were missing
        
        # 4. Complete counting logic with all sentiment cases
        if classification == NEU_LABEL:
            total_neutral += 1
        elif classification == POS_LABEL:
            total_positive += 1
        elif classification == NEG_LABEL:
            total_negative += 1
        else:
            # Catch unexpected labels to avoid silent failures
            print(f"Warning: Unknown sentiment label detected: {classification}")

# 5. Calculate and print final stats (matches your note about 90% neutral)
total_tweets = total_negative + total_positive + total_neutral
print(f"=== Sentiment Analysis Results ===")
print(f"Total tweets analyzed: {total_tweets}")
print(f"Neutral: {total_neutral} ({(total_neutral / total_tweets) * 100:.2f}%)")
print(f"Positive: {total_positive} ({(total_positive / total_tweets) * 100:.2f}%)")
print(f"Negative: {total_negative} ({(total_negative / total_tweets) * 100:.2f}%)")

Extra Notes for Your Use Case

  • JSON Data Handling: If your POStagged_tweets_lines.json has structured JSON objects per line (not just raw text), you'll need to parse each line first with json.loads() to extract the tweet text. For example:
    import json
    # ... inside the loop:
    tweet_obj = json.loads(line)
    cleaned_tweet = tweet_obj['text'].strip()  # Replace 'text' with your actual field name
    
  • Classifier Neutral Support: Since you mentioned 90% of your tags are neutral, double-check that your training data includes neutral examples. If your classifier wasn't trained on neutral sentiment, it won't output neu labels—this is crucial to getting accurate counts.

内容的提问来源于stack exchange,提问作者bybu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:26:34