You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求编写批量处理CSV文件的Python函数:筛选指定列并另存

Bulk CSV Processing: Extract Specific Columns & Save to Target Folder

Got it, let's fix this bulk CSV task so you don't have to manually process 500+ files one by one. Here's a clean, reliable solution that builds on the code you already have, with some key improvements to avoid headaches.

First, a quick note: Your original fileList only stores filenames, not full file paths—this would cause errors when trying to read files later (since Python won't know where to find them). Let's adjust the workflow to handle full paths directly, and wrap everything into a script that runs automatically.

Full Working Code

import os
import pandas as pd

# Set your source and target directories
source_dir = r'M:\BI\HisRms'
target_dir = r"C:\Users\jonathon.kindred\Desktop\RM\2019\FEB 2019"

# List of columns you want to keep (make sure these match your CSV headers exactly!)
required_columns = [
    'Purchase Order', 'SKU', 'Markdown', 'Landed Cost',
    'Original Price', 'Current Sale Price', 'Free Stock',
    'OPO', 'ID Style', 'Supplier Style No'
]

# Create target folder if it doesn't exist (no more manual folder creation!)
os.makedirs(target_dir, exist_ok=True)

# Loop through all CSV files in the source directory
for root, _, files in os.walk(source_dir):
    for filename in files:
        if filename.endswith('.csv'):
            # Build full paths for source and target files
            source_path = os.path.join(root, filename)
            target_path = os.path.join(target_dir, filename)
            
            try:
                # Read the CSV with your specified encoding and index column
                df = pd.read_csv(source_path, index_col=0, encoding='iso-8859-1')
                # Filter to keep only the columns we need
                filtered_df = df[required_columns]
                # Save to target folder without the index
                filtered_df.to_csv(target_path, index=False)
                print(f"✅ Processed: {filename}")
            except Exception as e:
                # Catch and print errors so the script doesn't crash mid-run
                print(f"❌ Failed to process {filename}: {str(e)}")

Key Details & Improvements

  • Full Path Handling: Uses os.path.join to build file paths, which works across Windows/macOS/Linux and avoids broken paths from manual slash typing.
  • Auto-Create Target Folder: os.makedirs(..., exist_ok=True) ensures your target directory exists—no need to create it manually before running the script.
  • Error Resilience: The try/except block catches issues like missing columns, corrupted files, or permission errors, prints a message, and keeps processing other files instead of crashing.
  • Progress Tracking: The print statements let you see which files are done and spot any failures at a glance.

Optional Optimizations

If you're dealing with extra-large CSV files (gigabytes in size), you can process them in chunks to save memory:

# Inside the loop, replace the read/filter/save steps with this:
chunk_size = 10000  # Adjust based on your memory
with pd.read_csv(source_path, index_col=0, encoding='iso-8859-1', chunksize=chunk_size) as reader:
    for i, chunk in enumerate(reader):
        filtered_chunk = chunk[required_columns]
        # Write header only for the first chunk
        filtered_chunk.to_csv(target_path, index=False, mode='a' if i > 0 else 'w', header=i == 0)

If you want to avoid overwriting existing files in the target folder, add this check before saving:

if not os.path.exists(target_path):
    filtered_df.to_csv(target_path, index=False)
else:
    print(f"⚠️ Skipping {filename} (already exists in target folder)")

内容的提问来源于stack exchange,提问作者JK1993

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 12:14:06