如何用Python从指定CSV文件中提取中性词至TXT文件
Hey there! As a Python newbie, file handling can feel a bit overwhelming at first, but let's break this down into simple, actionable steps to extract those neutral words from your CSV and save them to a TXT file. I'll give you two options—one using a popular library for easy CSV handling, and another using Python's built-in tools if you don't want to install anything extra.
Option 1: Use Pandas (Simpler for CSV Tasks)
Pandas is a go-to library for working with tabular data like CSVs, and it's super beginner-friendly once you set it up.
Step 1: Install Pandas
Open your terminal/command prompt and run this command to install the library:
pip install pandas
Step 2: Full Code with Explanations
Copy this code into a new Python file (e.g., extract_neutral_words.py), then adjust the file paths and threshold to match your needs:
import pandas as pd # 1. Load the CSV file (replace 'your_review_data.csv' with your actual file name/path) df = pd.read_csv('your_review_data.csv') # 2. Define what counts as a "neutral" word # We'll use words where the sentiment score's absolute value is ≤ 0.01 (adjust this number as needed!) neutral_words = df[df['Sentiment Score'].abs() <= 0.01]['Word'] # 3. Save the neutral words to a TXT file # index=False skips row numbers, header=False skips the "Word" column name neutral_words.to_csv('neutral_words.txt', index=False, header=False) # Optional: Print a sample to verify print("Sample neutral words:", neutral_words.head(10).tolist())
Option 2: Use Python's Built-in CSV Module (No External Libraries)
If you don't want to install pandas, you can use Python's built-in csv module. This is great for keeping things lightweight.
Full Code with Explanations
import csv neutral_words = [] # 1. Open and read the CSV file # Replace 'your_review_data.csv' with your actual file name/path with open('your_review_data.csv', 'r', newline='', encoding='utf-8') as csv_file: # Use DictReader to access columns by name reader = csv.DictReader(csv_file) # Loop through each row in the CSV for row in reader: # Convert the sentiment score from string to float sentiment_score = float(row['Sentiment Score']) # Check if the word is neutral (adjust the threshold here too!) if abs(sentiment_score) <= 0.01: neutral_words.append(row['Word']) # 2. Save the neutral words to a TXT file with open('neutral_words.txt', 'w', encoding='utf-8') as txt_file: # Write each word on a new line for word in neutral_words: txt_file.write(f"{word}\n") # Optional: Print a sample to verify print("Sample neutral words:", neutral_words[:10])
Key Notes for You:
- File Paths: If your CSV file isn't in the same folder as your Python script, replace the file name with the full path (e.g.,
'C:/Users/YourName/Documents/Amazon_Review_Sentiment_Analysis/review_data.csv'for Windows, or'/home/yourname/Documents/Amazon_Review_Sentiment_Analysis/review_data.csv'for macOS/Linux). - Neutral Threshold: The
0.01value is a starting point. If you want more words to count as neutral, increase it (e.g.,0.02); if you want stricter neutrality, decrease it (e.g.,0.005). - Encoding: Using
encoding='utf-8'ensures you don't get weird character issues when reading/writing files.
内容的提问来源于stack exchange,提问作者ANiK3T

