请求编写批量处理同文件夹下400个CSV文件的Python脚本
Batch Processing CSV Files with Unique Outputs in Python
Got it, let's walk through how to automate this batch processing task properly. I'll provide a complete, customizable script that handles all the steps—from finding your CSV files to saving unique outputs in a dedicated folder.
Step-by-Step Explanation & Script
First, here's the full script. I'll break down each part after so you can adapt it to your specific processing needs:
import os import glob # If you're using pandas for processing (common for CSV tasks), uncomment this: # import pandas as pd # If using the built-in csv module, uncomment this: # import csv def process_single_csv(input_file): """ Replace this function with your actual CSV processing logic. Args: input_file (str): Path to the input CSV file. Returns: Processed data (e.g., pandas DataFrame, list of lists, etc.) """ # Example processing using pandas (replace with your code): # df = pd.read_csv(input_file) # # Add your processing steps here: clean data, compute metrics, filter rows, etc. # processed_df = df # Placeholder for your processed data # return processed_df # Example processing using built-in csv module: # processed_data = [] # with open(input_file, 'r', newline='', encoding='utf-8') as infile: # reader = csv.reader(infile) # # Process each row as needed (e.g., modify values, add columns) # for row in reader: # processed_row = row # Placeholder for your row processing # processed_data.append(processed_row) # return processed_data # REMINDER: Delete the placeholder comments above and add your actual processing code! raise NotImplementedError("Please replace this function with your CSV processing logic.") def save_processed_data(processed_data, output_file_path): """ Save processed data to a CSV file. Adapt this based on your processing output type. Args: processed_data: Data from process_single_csv() output_file_path (str): Path to save the output CSV. """ # If using pandas: # processed_data.to_csv(output_file_path, index=False) # If using csv module: # with open(output_file_path, 'w', newline='', encoding='utf-8') as outfile: # writer = csv.writer(outfile) # writer.writerows(processed_data) # REMINDER: Uncomment the appropriate section based on your processing method! raise NotImplementedError("Please adapt this function to save your processed data.") if __name__ == "__main__": # Configuration - adjust these paths to match your setup INPUT_FOLDER = "." # Use "." for the same folder as the script, or replace with a full path like "C:/your/csv/folder" OUTPUT_FOLDER = "processed_csv_results" # Create output folder if it doesn't exist (no error if it already exists) os.makedirs(OUTPUT_FOLDER, exist_ok=True) # Get all CSV files in the input folder (avoids non-CSV files) csv_files = glob.glob(os.path.join(INPUT_FOLDER, "*.csv")) if not csv_files: print("No CSV files found in the input folder!") exit() # Process each CSV file one by one for csv_path in csv_files: # Extract original filename (without extension) to make a unique output name file_name = os.path.splitext(os.path.basename(csv_path))[0] output_file_name = f"{file_name}_processed.csv" output_file_path = os.path.join(OUTPUT_FOLDER, output_file_name) try: print(f"Processing {csv_path}...") processed_data = process_single_csv(csv_path) save_processed_data(processed_data, output_file_path) print(f"Successfully saved output to {output_file_path}") except Exception as e: print(f"Error processing {csv_path}: {str(e)}") print("Batch processing complete!")
Key Details to Adapt
- Configuration: Update
INPUT_FOLDERto point to where your 400 CSVs are stored (use"."if they're in the same folder as the script).OUTPUT_FOLDERis the name of the dedicated results folder—it will be created automatically if it doesn't exist. - Processing Logic: Replace the placeholder code in
process_single_csv()with your actual steps (cleaning data, calculating metrics, etc.). The script supports both pandas (most common for CSV tasks) and the built-incsvmodule. - Saving Logic: Adjust
save_processed_data()to match your processing method. Uncomment the pandas section if you're using that library, or thecsvmodule section if you prefer the built-in tool. - Unique Filenames: The script appends
_processedto the original filename (e.g.,sales_data.csvbecomessales_data_processed.csv) to ensure every output has a unique, recognizable name.
Why This Works
glob.glob()safely targets only CSV files, so you won't accidentally process other types of files in the folder.os.makedirs(..., exist_ok=True)handles folder creation without throwing errors if the folder already exists.- The
try/exceptblock ensures that if one file fails to process, the script keeps running for the remaining files instead of crashing entirely.
内容的提问来源于stack exchange,提问作者Pronomita Dey
相关产品推荐
相关产品推荐

