请求编写批量处理CSV文件的Python函数:筛选指定列并另存
Got it, let's fix this bulk CSV task so you don't have to manually process 500+ files one by one. Here's a clean, reliable solution that builds on the code you already have, with some key improvements to avoid headaches.
First, a quick note: Your original fileList only stores filenames, not full file paths—this would cause errors when trying to read files later (since Python won't know where to find them). Let's adjust the workflow to handle full paths directly, and wrap everything into a script that runs automatically.
Full Working Code
import os import pandas as pd # Set your source and target directories source_dir = r'M:\BI\HisRms' target_dir = r"C:\Users\jonathon.kindred\Desktop\RM\2019\FEB 2019" # List of columns you want to keep (make sure these match your CSV headers exactly!) required_columns = [ 'Purchase Order', 'SKU', 'Markdown', 'Landed Cost', 'Original Price', 'Current Sale Price', 'Free Stock', 'OPO', 'ID Style', 'Supplier Style No' ] # Create target folder if it doesn't exist (no more manual folder creation!) os.makedirs(target_dir, exist_ok=True) # Loop through all CSV files in the source directory for root, _, files in os.walk(source_dir): for filename in files: if filename.endswith('.csv'): # Build full paths for source and target files source_path = os.path.join(root, filename) target_path = os.path.join(target_dir, filename) try: # Read the CSV with your specified encoding and index column df = pd.read_csv(source_path, index_col=0, encoding='iso-8859-1') # Filter to keep only the columns we need filtered_df = df[required_columns] # Save to target folder without the index filtered_df.to_csv(target_path, index=False) print(f"✅ Processed: {filename}") except Exception as e: # Catch and print errors so the script doesn't crash mid-run print(f"❌ Failed to process {filename}: {str(e)}")
Key Details & Improvements
- Full Path Handling: Uses
os.path.jointo build file paths, which works across Windows/macOS/Linux and avoids broken paths from manual slash typing. - Auto-Create Target Folder:
os.makedirs(..., exist_ok=True)ensures your target directory exists—no need to create it manually before running the script. - Error Resilience: The
try/exceptblock catches issues like missing columns, corrupted files, or permission errors, prints a message, and keeps processing other files instead of crashing. - Progress Tracking: The print statements let you see which files are done and spot any failures at a glance.
Optional Optimizations
If you're dealing with extra-large CSV files (gigabytes in size), you can process them in chunks to save memory:
# Inside the loop, replace the read/filter/save steps with this: chunk_size = 10000 # Adjust based on your memory with pd.read_csv(source_path, index_col=0, encoding='iso-8859-1', chunksize=chunk_size) as reader: for i, chunk in enumerate(reader): filtered_chunk = chunk[required_columns] # Write header only for the first chunk filtered_chunk.to_csv(target_path, index=False, mode='a' if i > 0 else 'w', header=i == 0)
If you want to avoid overwriting existing files in the target folder, add this check before saving:
if not os.path.exists(target_path): filtered_df.to_csv(target_path, index=False) else: print(f"⚠️ Skipping {filename} (already exists in target folder)")
内容的提问来源于stack exchange,提问作者JK1993

