无需Pandas:Python导入多文件夹CSV并添加文件名列的实现
Got it, I totally get why you want to skip Pandas here—processing hundreds of CSVs with it can get sluggish. Your original binary copy approach is fast but doesn’t handle adding a filename column or duplicate headers. Let’s fix that while keeping the operation efficient:
Modified Basic Version (Text Mode)
This version reads each CSV as text, adds the filename column, and ensures we only write the header once:
import glob import os from os.path import basename # Set your target directory and output path input_dir = "D:\\DATA" output_file = "D:\\AllData.csv" os.chdir(input_dir) all_files = glob.glob("**/*.csv", recursive=True) with open(output_file, 'w', encoding='utf-8', newline='') as outfile: header_written = False for fname in all_files: # Get the filename (use fname if you want full path instead of just the file name) source_filename = basename(fname) with open(fname, 'r', encoding='utf-8', newline='') as infile: reader = infile.readlines() # Handle header if not header_written: # Add new column name to header header = reader[0].strip() + ',source_filename\n' outfile.write(header) header_written = True # Write the first row's data with filename if len(reader) > 1: for line in reader[1:]: outfile.write(f"{line.strip()},{source_filename}\n") else: # Skip header for subsequent files for line in reader[1:]: outfile.write(f"{line.strip()},{source_filename}\n")
Key Improvements:
- Avoids duplicate headers: Only writes the header once at the start
- Adds filename column: Appends
source_filename(or full path if you usefnameinstead ofbasename(fname)) to every row - Text mode with encoding: Uses
utf-8encoding to handle special characters, andnewline=''to handle cross-platform line endings correctly - Keeps speed: Reads files efficiently line by line, no heavy libraries involved
Robust Version (Using csv Module)
If your CSVs contain commas in fields (e.g., quoted values like "Doe, John"), the basic string concatenation will break column alignment. For these cases, use the built-in csv module—it handles quoted fields safely:
import glob import os import csv from os.path import basename input_dir = "D:\\DATA" output_file = "D:\\AllData.csv" os.chdir(input_dir) all_files = glob.glob("**/*.csv", recursive=True) with open(output_file, 'w', encoding='utf-8', newline='') as outfile: writer = None for fname in all_files: source_filename = basename(fname) with open(fname, 'r', encoding='utf-8', newline='') as infile: reader = csv.reader(infile) if writer is None: # Create writer with added column name header = next(reader) header.append('source_filename') writer = csv.writer(outfile) writer.writerow(header) # Write all rows with filename appended for row in reader: row.append(source_filename) writer.writerow(row)
This version is slightly more CPU-intensive but far safer for real-world CSV data with special formatting.
内容的提问来源于stack exchange,提问作者user9853666

