You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求编写批量处理同文件夹下400个CSV文件的Python脚本

Batch Processing CSV Files with Unique Outputs in Python

Got it, let's walk through how to automate this batch processing task properly. I'll provide a complete, customizable script that handles all the steps—from finding your CSV files to saving unique outputs in a dedicated folder.

Step-by-Step Explanation & Script

First, here's the full script. I'll break down each part after so you can adapt it to your specific processing needs:

import os
import glob
# If you're using pandas for processing (common for CSV tasks), uncomment this:
# import pandas as pd
# If using the built-in csv module, uncomment this:
# import csv

def process_single_csv(input_file):
    """
    Replace this function with your actual CSV processing logic.
    Args:
        input_file (str): Path to the input CSV file.
    Returns:
        Processed data (e.g., pandas DataFrame, list of lists, etc.)
    """
    # Example processing using pandas (replace with your code):
    # df = pd.read_csv(input_file)
    # # Add your processing steps here: clean data, compute metrics, filter rows, etc.
    # processed_df = df  # Placeholder for your processed data
    # return processed_df

    # Example processing using built-in csv module:
    # processed_data = []
    # with open(input_file, 'r', newline='', encoding='utf-8') as infile:
    #     reader = csv.reader(infile)
    #     # Process each row as needed (e.g., modify values, add columns)
    #     for row in reader:
    #         processed_row = row  # Placeholder for your row processing
    #         processed_data.append(processed_row)
    # return processed_data

    # REMINDER: Delete the placeholder comments above and add your actual processing code!
    raise NotImplementedError("Please replace this function with your CSV processing logic.")

def save_processed_data(processed_data, output_file_path):
    """
    Save processed data to a CSV file. Adapt this based on your processing output type.
    Args:
        processed_data: Data from process_single_csv()
        output_file_path (str): Path to save the output CSV.
    """
    # If using pandas:
    # processed_data.to_csv(output_file_path, index=False)

    # If using csv module:
    # with open(output_file_path, 'w', newline='', encoding='utf-8') as outfile:
    #     writer = csv.writer(outfile)
    #     writer.writerows(processed_data)

    # REMINDER: Uncomment the appropriate section based on your processing method!
    raise NotImplementedError("Please adapt this function to save your processed data.")

if __name__ == "__main__":
    # Configuration - adjust these paths to match your setup
    INPUT_FOLDER = "."  # Use "." for the same folder as the script, or replace with a full path like "C:/your/csv/folder"
    OUTPUT_FOLDER = "processed_csv_results"

    # Create output folder if it doesn't exist (no error if it already exists)
    os.makedirs(OUTPUT_FOLDER, exist_ok=True)

    # Get all CSV files in the input folder (avoids non-CSV files)
    csv_files = glob.glob(os.path.join(INPUT_FOLDER, "*.csv"))

    if not csv_files:
        print("No CSV files found in the input folder!")
        exit()

    # Process each CSV file one by one
    for csv_path in csv_files:
        # Extract original filename (without extension) to make a unique output name
        file_name = os.path.splitext(os.path.basename(csv_path))[0]
        output_file_name = f"{file_name}_processed.csv"
        output_file_path = os.path.join(OUTPUT_FOLDER, output_file_name)

        try:
            print(f"Processing {csv_path}...")
            processed_data = process_single_csv(csv_path)
            save_processed_data(processed_data, output_file_path)
            print(f"Successfully saved output to {output_file_path}")
        except Exception as e:
            print(f"Error processing {csv_path}: {str(e)}")

    print("Batch processing complete!")

Key Details to Adapt

  • Configuration: Update INPUT_FOLDER to point to where your 400 CSVs are stored (use "." if they're in the same folder as the script). OUTPUT_FOLDER is the name of the dedicated results folder—it will be created automatically if it doesn't exist.
  • Processing Logic: Replace the placeholder code in process_single_csv() with your actual steps (cleaning data, calculating metrics, etc.). The script supports both pandas (most common for CSV tasks) and the built-in csv module.
  • Saving Logic: Adjust save_processed_data() to match your processing method. Uncomment the pandas section if you're using that library, or the csv module section if you prefer the built-in tool.
  • Unique Filenames: The script appends _processed to the original filename (e.g., sales_data.csv becomes sales_data_processed.csv) to ensure every output has a unique, recognizable name.

Why This Works

  • glob.glob() safely targets only CSV files, so you won't accidentally process other types of files in the folder.
  • os.makedirs(..., exist_ok=True) handles folder creation without throwing errors if the folder already exists.
  • The try/except block ensures that if one file fails to process, the script keeps running for the remaining files instead of crashing entirely.

内容的提问来源于stack exchange,提问作者Pronomita Dey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:14:41