如何用Python遍历目录中TXT垃圾邮件文件,提取特征并输出至CSV
Alright, let's build a robust Python solution that handles your spam email processing task exactly as described. I'll break this down into actionable code and explain each part so you can adapt it to your needs.
Step 1: Setup Dependencies
We'll use Python's built-in libraries, so no extra installs are needed. We'll rely on os for directory traversal and csv for writing the output file.
Step 2: Complete Code Implementation
Copy this code into a Python file (e.g., spam_processor.py):
import os import csv # Configuration - adjust these paths and settings to match your environment SPAM_DIRECTORY = "./spam_emails" # Path to your folder of spam .txt files OUTPUT_CSV = "./spam_features.csv" # Path where the results will be saved MALICIOUS_WORDS = ["free money", "win instantly", "click to claim", "verify your account", "spam"] # Customize your malicious terms def process_single_email(file_path, malicious_terms): """Process a single email file and return feature results.""" try: # Read the email content, handle encoding issues gracefully with open(file_path, 'r', encoding='utf-8', errors='ignore') as email_file: email_content = email_file.read().lower() # Convert to lowercase for case-insensitive checks # Check if any malicious word exists in the content contains_malicious = any(term in email_content for term in malicious_terms) return 'TRUE' if contains_malicious else 'FALSE' except Exception as e: # Log errors but don't stop the entire process print(f"Warning: Failed to process {file_path} - {str(e)}") return 'FALSE' # Default to FALSE if processing fails def main(): # Create output directory if it doesn't exist os.makedirs(os.path.dirname(OUTPUT_CSV), exist_ok=True) # Get all .txt files in the spam directory spam_files = [filename for filename in os.listdir(SPAM_DIRECTORY) if filename.endswith('.txt')] if not spam_files: print("No .txt files found in the specified directory.") return # Write results to CSV with open(OUTPUT_CSV, 'w', newline='', encoding='utf-8') as csv_file: csv_writer = csv.writer(csv_file) # Write header row csv_writer.writerow(['Email_Filename', 'Contains_Malicious_Word']) # Process each file and write to CSV for filename in spam_files: full_file_path = os.path.join(SPAM_DIRECTORY, filename) feature_result = process_single_email(full_file_path, MALICIOUS_WORDS) csv_writer.writerow([filename, feature_result]) print(f"Processing finished! {len(spam_files)} emails processed. Results saved to {OUTPUT_CSV}") if __name__ == "__main__": main()
Step 3: Key Code Explanations
- Directory Traversal: The code uses
os.listdirto filter only.txtfiles, ensuring we don't process unrelated files. - Robust File Reading: We use
encoding='utf-8'witherrors='ignore'to handle any weird encoding issues in spam emails (common in real-world datasets). Converting content to lowercase ensures our malicious word checks aren't case-sensitive. - Feature Extraction: The core check looks for any of your defined malicious terms in the email content. It returns
TRUEorFALSEexactly as required. - CSV Output: We write a clear header row followed by each email's filename and feature result. The
newline=''argument prevents extra blank lines in the CSV on Windows.
Step 4: Extending the Solution
This code is easy to expand for more features. For example, if you want to add a "Contains_Link" feature:
- Add a check in
process_single_email(e.g., look forhttp://orhttps://in the content) - Update the header row to include
'Contains_Link' - Modify the
csv_writer.writerowline to include the new feature result
For large datasets (thousands of emails), you can speed up processing using Python's concurrent.futures module to run file processing in parallel.
内容的提问来源于stack exchange,提问作者HendoRelish

