计算文件夹中文本文件复合情感值平均值时遇技术问题
Hey there! Let's figure out how to calculate the average compound sentiment score for each text file in your folder—building right off the code you already have for sentence-level analysis. Here's a step-by-step solution with fully working code and explanations:
Fixing Average Compound Sentiment Calculation for Folder Text Files
Full Working Code
import glob import os import nltk.data from nltk.sentiment.vader import SentimentIntensityAnalyzer from nltk import sentiment from nltk import word_tokenize # Initialize VADER and sentence tokenizer (your original setup!) sid = SentimentIntensityAnalyzer() tokenizer = nltk.data.load('tokenizers/punkt/english.pickle') # Replace this with your actual folder path containing text files folder_path = "/path/to/your/text/files" # Loop through every .txt file in the target folder for file_path in glob.glob(os.path.join(folder_path, "*.txt")): # Grab the filename to keep track of which file we're processing filename = os.path.basename(file_path) # Read the full content of the text file with open(file_path, 'r', encoding='utf-8') as file: text_content = file.read() # Split the text into individual sentences (using your tokenizer) sentences = tokenizer.tokenize(text_content) # Collect compound sentiment scores for all sentences in the file compound_scores = [] for sentence in sentences: # Get VADER's sentiment scores for the sentence sentiment_scores = sid.polarity_scores(sentence) # Extract the compound score (the overall sentiment value: -1 to 1) compound_scores.append(sentiment_scores['compound']) # Calculate and print the average compound score for the file if compound_scores: average_compound = sum(compound_scores) / len(compound_scores) print(f"✅ File: {filename} | Average Compound Sentiment: {average_compound:.4f}") else: print(f"⚠️ File: {filename} | No valid sentences found (empty or unreadable file?)")
Key Details & Troubleshooting Tips
- Folder Path Setup: Make sure to replace
/path/to/your/text/fileswith the actual path to your folder. Usingos.path.joinensures this works across Windows, Mac, and Linux without path separator issues. - Encoding Fixes: If you get encoding errors when reading files, try swapping
encoding='utf-8'forencoding='latin-1'or addingerrors='ignore'to theopen()function. - Empty File Handling: The check for
compound_scoresprevents division-by-zero errors if a file is empty or has no recognizable sentences. - Optional: Save Results to CSV: If you want to store the results instead of just printing them, add this code (place it right before the file loop):
import csv # Create a CSV file to save results with open("sentiment_results.csv", 'w', newline='', encoding='utf-8') as csvfile: writer = csv.writer(csvfile) writer.writerow(["Filename", "Average Compound Sentiment"]) # Then inside your file loop, replace the print statement with: if compound_scores: writer.writerow([filename, round(average_compound, 4)]) else: writer.writerow([filename, "N/A"])
内容的提问来源于stack exchange,提问作者K. SOT
相关产品推荐
相关产品推荐

