You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让Python代码自动遍历91个公司文件夹执行Twitter情感分析

自动遍历多文件夹执行Twitter推文情感分析

问题描述

我桌面上有一个companyfollowerstweets文件夹,里面包含91个以followerstweets(公司名称)命名的子文件夹。每个子文件夹里有200个CSV文件,每个文件存储对应公司一位Twitter粉丝的最新推文。我需要对每个公司的200位粉丝各取前200条推文做情感分析,最终得到每个公司总计40000条推文的正负占比、平均置信度等结果。目前我的代码只能手动指定单个公司名称来遍历文件夹,希望修改成自动遍历全部91个文件夹。

现有代码(仅处理单个公司):

import nltk
import csv
import sklearn
import nltk, string, numpy
from sklearn.feature_extraction.text import TfidfVectorizer
from collections import defaultdict
from sklearn.feature_extraction.text import CountVectorizer
columns = defaultdict(list)
from nltk.corpus import stopwords
from sklearn.feature_extraction.text import TfidfTransformer
import math
import sentiment_mod as s
import glob
import itertools
lijst = glob.glob('companyfollowerstweets/followerstweetsCisco/*.csv')
tweets1 = []
sent1 = []
print(lijst[0])
for item in lijst:
    stopwords_set = set(stopwords.words("english"))
    with open(item, encoding = 'latin-1') as d:
        reader1=csv.reader(d)
        next(reader1)
        for row in itertools.islice(reader1,200):
            tweets1.extend([row[2]])
    words_cleaned = [" ".join([words for words in sentence.split() if 'http' not in words and not words.startswith('@')]) for sentence in tweets1]
    words_filtered = [e.lower() for e in words_cleaned]
    words_without_stopwords = [word for word in words_filtered if not word in stopwords_set]
    tweets1 = words_without_stopwords
    tweets1 = list(filter(None, tweets1))
for d in tweets1:
    new1 = s.sentiment(d)
    sent1.extend(new1)
total1 = len(sent1)/2
neg_percentage1 = (sent1.count("neg")/total1)*100
pos_percentage1 = (sent1.count("pos")/total1)*100
res = sum(sent1[1::2])/total1
low = min(sent1[1::2])
high = max(sent1[1::2])
print("% of negative Tweets:", neg_percentage1)
print("% of positive Tweets:", pos_percentage1)
print("Total number of Tweets:", total1)
print("Average confidence:", res)
print("min confidence:", low)
print("max confidence:", high)

修改方案与代码

我帮你调整了代码,核心是增加外层循环遍历所有公司文件夹,同时优化了变量作用域和数据清洗逻辑,确保每个公司的分析结果独立且准确:

import nltk
import csv
import sklearn
import string
import numpy
import os
from sklearn.feature_extraction.text import TfidfVectorizer, CountVectorizer, TfidfTransformer
from collections import defaultdict
from nltk.corpus import stopwords
import math
import sentiment_mod as s
import glob
import itertools

# 提前下载停用词(首次运行需要,之后可注释)
nltk.download('stopwords')
stopwords_set = set(stopwords.words("english"))

# 获取所有公司的子文件夹路径
company_folders = glob.glob('companyfollowerstweets/followerstweets*')

# 遍历每个公司文件夹
for folder_path in company_folders:
    # 从文件夹名提取公司名称(比如"followerstweetsCisco" -> "Cisco")
    company_name = os.path.basename(folder_path).replace('followerstweets', '')
    print(f"\n=== 正在处理公司: {company_name} ===")
    
    # 重置当前公司的推文和情感结果列表(避免不同公司数据混淆)
    tweets1 = []
    sent1 = []
    
    # 获取当前文件夹下的所有CSV文件
    csv_files = glob.glob(f'{folder_path}/*.csv')
    print(f"找到 {len(csv_files)} 个粉丝推文文件")
    
    # 遍历当前公司的所有CSV文件
    for item in csv_files:
        with open(item, encoding='latin-1') as d:
            reader1 = csv.reader(d)
            next(reader1)  # 跳过CSV表头
            
            # 提取当前粉丝的前200条推文
            for row in itertools.islice(reader1, 200):
                tweet_text = row[2]
                # 实时清洗推文:移除链接、@提及,转小写,过滤停用词
                cleaned_words = [word for word in tweet_text.split() if 'http' not in word and not word.startswith('@')]
                cleaned_tweet = " ".join(cleaned_words).lower()
                # 过滤空字符串和纯停用词的推文
                if cleaned_tweet and cleaned_tweet not in stopwords_set:
                    tweets1.append(cleaned_tweet)
    
    # 对当前公司的所有有效推文执行情感分析
    for tweet in tweets1:
        sentiment_result = s.sentiment(tweet)
        sent1.extend(sentiment_result)
    
    # 计算并输出统计结果(避免除以零错误)
    total_tweets = len(sent1) // 2  # 每个结果是(标签,置信度),所以长度是2倍推文数
    if total_tweets == 0:
        print("⚠️ 该公司没有有效推文可分析")
        continue
    
    neg_count = sent1.count("neg")
    pos_count = sent1.count("pos")
    neg_percentage = (neg_count / total_tweets) * 100
    pos_percentage = (pos_count / total_tweets) * 100
    
    confidences = sent1[1::2]
    avg_confidence = sum(confidences) / total_tweets
    min_confidence = min(confidences)
    max_confidence = max(confidences)
    
    # 格式化输出结果
    print(f"📊 公司: {company_name}")
    print(f"负面推文占比: {neg_percentage:.2f}%")
    print(f"正面推文占比: {pos_percentage:.2f}%")
    print(f"总有效推文数: {total_tweets}")
    print(f"平均置信度: {avg_confidence:.4f}")
    print(f"最低置信度: {min_confidence:.4f}")
    print(f"最高置信度: {max_confidence:.4f}")

关键改动说明

  • 自动遍历所有公司:用glob.glob('companyfollowerstweets/followerstweets*')匹配所有公司子文件夹,无需手动输入公司名称。
  • 隔离公司数据:每次处理新公司时重置tweets1和sent1,确保不同公司的分析结果不会互相干扰。
  • 优化清洗逻辑:把推文清洗移到读取CSV的循环中,避免重复处理整个列表,提升运行效率。
  • 增加容错处理:如果某个公司没有有效推文,会跳过统计步骤,避免出现除以零的错误。
  • 清晰的结果标识:从文件夹名提取公司名称,输出时明确标注每个公司的结果,方便查看。

内容的提问来源于stack exchange,提问作者Nienke Luirink

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:57:39