如何让Python代码自动遍历91个公司文件夹执行Twitter情感分析
自动遍历多文件夹执行Twitter推文情感分析
问题描述
我桌面上有一个companyfollowerstweets文件夹,里面包含91个以followerstweets(公司名称)命名的子文件夹。每个子文件夹里有200个CSV文件,每个文件存储对应公司一位Twitter粉丝的最新推文。我需要对每个公司的200位粉丝各取前200条推文做情感分析,最终得到每个公司总计40000条推文的正负占比、平均置信度等结果。目前我的代码只能手动指定单个公司名称来遍历文件夹,希望修改成自动遍历全部91个文件夹。
现有代码(仅处理单个公司):
import nltk import csv import sklearn import nltk, string, numpy from sklearn.feature_extraction.text import TfidfVectorizer from collections import defaultdict from sklearn.feature_extraction.text import CountVectorizer columns = defaultdict(list) from nltk.corpus import stopwords from sklearn.feature_extraction.text import TfidfTransformer import math import sentiment_mod as s import glob import itertools lijst = glob.glob('companyfollowerstweets/followerstweetsCisco/*.csv') tweets1 = [] sent1 = [] print(lijst[0]) for item in lijst: stopwords_set = set(stopwords.words("english")) with open(item, encoding = 'latin-1') as d: reader1=csv.reader(d) next(reader1) for row in itertools.islice(reader1,200): tweets1.extend([row[2]]) words_cleaned = [" ".join([words for words in sentence.split() if 'http' not in words and not words.startswith('@')]) for sentence in tweets1] words_filtered = [e.lower() for e in words_cleaned] words_without_stopwords = [word for word in words_filtered if not word in stopwords_set] tweets1 = words_without_stopwords tweets1 = list(filter(None, tweets1)) for d in tweets1: new1 = s.sentiment(d) sent1.extend(new1) total1 = len(sent1)/2 neg_percentage1 = (sent1.count("neg")/total1)*100 pos_percentage1 = (sent1.count("pos")/total1)*100 res = sum(sent1[1::2])/total1 low = min(sent1[1::2]) high = max(sent1[1::2]) print("% of negative Tweets:", neg_percentage1) print("% of positive Tweets:", pos_percentage1) print("Total number of Tweets:", total1) print("Average confidence:", res) print("min confidence:", low) print("max confidence:", high)
修改方案与代码
我帮你调整了代码,核心是增加外层循环遍历所有公司文件夹,同时优化了变量作用域和数据清洗逻辑,确保每个公司的分析结果独立且准确:
import nltk import csv import sklearn import string import numpy import os from sklearn.feature_extraction.text import TfidfVectorizer, CountVectorizer, TfidfTransformer from collections import defaultdict from nltk.corpus import stopwords import math import sentiment_mod as s import glob import itertools # 提前下载停用词(首次运行需要,之后可注释) nltk.download('stopwords') stopwords_set = set(stopwords.words("english")) # 获取所有公司的子文件夹路径 company_folders = glob.glob('companyfollowerstweets/followerstweets*') # 遍历每个公司文件夹 for folder_path in company_folders: # 从文件夹名提取公司名称(比如"followerstweetsCisco" -> "Cisco") company_name = os.path.basename(folder_path).replace('followerstweets', '') print(f"\n=== 正在处理公司: {company_name} ===") # 重置当前公司的推文和情感结果列表(避免不同公司数据混淆) tweets1 = [] sent1 = [] # 获取当前文件夹下的所有CSV文件 csv_files = glob.glob(f'{folder_path}/*.csv') print(f"找到 {len(csv_files)} 个粉丝推文文件") # 遍历当前公司的所有CSV文件 for item in csv_files: with open(item, encoding='latin-1') as d: reader1 = csv.reader(d) next(reader1) # 跳过CSV表头 # 提取当前粉丝的前200条推文 for row in itertools.islice(reader1, 200): tweet_text = row[2] # 实时清洗推文:移除链接、@提及,转小写,过滤停用词 cleaned_words = [word for word in tweet_text.split() if 'http' not in word and not word.startswith('@')] cleaned_tweet = " ".join(cleaned_words).lower() # 过滤空字符串和纯停用词的推文 if cleaned_tweet and cleaned_tweet not in stopwords_set: tweets1.append(cleaned_tweet) # 对当前公司的所有有效推文执行情感分析 for tweet in tweets1: sentiment_result = s.sentiment(tweet) sent1.extend(sentiment_result) # 计算并输出统计结果(避免除以零错误) total_tweets = len(sent1) // 2 # 每个结果是(标签,置信度),所以长度是2倍推文数 if total_tweets == 0: print("⚠️ 该公司没有有效推文可分析") continue neg_count = sent1.count("neg") pos_count = sent1.count("pos") neg_percentage = (neg_count / total_tweets) * 100 pos_percentage = (pos_count / total_tweets) * 100 confidences = sent1[1::2] avg_confidence = sum(confidences) / total_tweets min_confidence = min(confidences) max_confidence = max(confidences) # 格式化输出结果 print(f"📊 公司: {company_name}") print(f"负面推文占比: {neg_percentage:.2f}%") print(f"正面推文占比: {pos_percentage:.2f}%") print(f"总有效推文数: {total_tweets}") print(f"平均置信度: {avg_confidence:.4f}") print(f"最低置信度: {min_confidence:.4f}") print(f"最高置信度: {max_confidence:.4f}")
关键改动说明
- 自动遍历所有公司:用
glob.glob('companyfollowerstweets/followerstweets*')匹配所有公司子文件夹,无需手动输入公司名称。 - 隔离公司数据:每次处理新公司时重置
tweets1和sent1,确保不同公司的分析结果不会互相干扰。 - 优化清洗逻辑:把推文清洗移到读取CSV的循环中,避免重复处理整个列表,提升运行效率。
- 增加容错处理:如果某个公司没有有效推文,会跳过统计步骤,避免出现除以零的错误。
- 清晰的结果标识:从文件夹名提取公司名称,输出时明确标注每个公司的结果,方便查看。
内容的提问来源于stack exchange,提问作者Nienke Luirink
相关产品推荐
相关产品推荐

