You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python遍历目录中TXT垃圾邮件文件,提取特征并输出至CSV

Alright, let's build a robust Python solution that handles your spam email processing task exactly as described. I'll break this down into actionable code and explain each part so you can adapt it to your needs.

Python Solution for Spam Email Feature Extraction and CSV Export

Step 1: Setup Dependencies

We'll use Python's built-in libraries, so no extra installs are needed. We'll rely on os for directory traversal and csv for writing the output file.

Step 2: Complete Code Implementation

Copy this code into a Python file (e.g., spam_processor.py):

import os
import csv

# Configuration - adjust these paths and settings to match your environment
SPAM_DIRECTORY = "./spam_emails"  # Path to your folder of spam .txt files
OUTPUT_CSV = "./spam_features.csv"  # Path where the results will be saved
MALICIOUS_WORDS = ["free money", "win instantly", "click to claim", "verify your account", "spam"]  # Customize your malicious terms

def process_single_email(file_path, malicious_terms):
    """Process a single email file and return feature results."""
    try:
        # Read the email content, handle encoding issues gracefully
        with open(file_path, 'r', encoding='utf-8', errors='ignore') as email_file:
            email_content = email_file.read().lower()  # Convert to lowercase for case-insensitive checks
        
        # Check if any malicious word exists in the content
        contains_malicious = any(term in email_content for term in malicious_terms)
        return 'TRUE' if contains_malicious else 'FALSE'
    
    except Exception as e:
        # Log errors but don't stop the entire process
        print(f"Warning: Failed to process {file_path} - {str(e)}")
        return 'FALSE'  # Default to FALSE if processing fails

def main():
    # Create output directory if it doesn't exist
    os.makedirs(os.path.dirname(OUTPUT_CSV), exist_ok=True)
    
    # Get all .txt files in the spam directory
    spam_files = [filename for filename in os.listdir(SPAM_DIRECTORY) if filename.endswith('.txt')]
    
    if not spam_files:
        print("No .txt files found in the specified directory.")
        return
    
    # Write results to CSV
    with open(OUTPUT_CSV, 'w', newline='', encoding='utf-8') as csv_file:
        csv_writer = csv.writer(csv_file)
        # Write header row
        csv_writer.writerow(['Email_Filename', 'Contains_Malicious_Word'])
        
        # Process each file and write to CSV
        for filename in spam_files:
            full_file_path = os.path.join(SPAM_DIRECTORY, filename)
            feature_result = process_single_email(full_file_path, MALICIOUS_WORDS)
            csv_writer.writerow([filename, feature_result])
    
    print(f"Processing finished! {len(spam_files)} emails processed. Results saved to {OUTPUT_CSV}")

if __name__ == "__main__":
    main()

Step 3: Key Code Explanations

  • Directory Traversal: The code uses os.listdir to filter only .txt files, ensuring we don't process unrelated files.
  • Robust File Reading: We use encoding='utf-8' with errors='ignore' to handle any weird encoding issues in spam emails (common in real-world datasets). Converting content to lowercase ensures our malicious word checks aren't case-sensitive.
  • Feature Extraction: The core check looks for any of your defined malicious terms in the email content. It returns TRUE or FALSE exactly as required.
  • CSV Output: We write a clear header row followed by each email's filename and feature result. The newline='' argument prevents extra blank lines in the CSV on Windows.

Step 4: Extending the Solution

This code is easy to expand for more features. For example, if you want to add a "Contains_Link" feature:

  1. Add a check in process_single_email (e.g., look for http:// or https:// in the content)
  2. Update the header row to include 'Contains_Link'
  3. Modify the csv_writer.writerow line to include the new feature result

For large datasets (thousands of emails), you can speed up processing using Python's concurrent.futures module to run file processing in parallel.


内容的提问来源于stack exchange,提问作者HendoRelish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:54:23