如何用Python批量转换指定命名模式的TSV为CSV并合并
Got it, let's work through this batch processing task efficiently! You already know how to handle a single TSV file, so scaling this up to process all your files, convert them to CSV, and merge into one big dataset is totally doable with a few tweaks.
Step 1: Import Required Libraries
First, we'll need these core libraries to handle file matching, data reading, and conversion:
import pandas as pd import glob import os
Step 2: Set Up File Path & Match Target Files
Use glob to automatically find all TSV files that follow your CustomerData_YYYYMMDDXXXXXX.tsv naming pattern. This is way more efficient than manually listing files:
# Replace this with your actual folder path folder_path = r"C:\Users\username\Desktop\folder" # Grab all TSV files matching the naming pattern tsv_files = glob.glob(os.path.join(folder_path, "CustomerData_*.tsv"))
Step 3: Batch Process Each File
Loop through each matched file, read it (using your existing single-file logic), convert it to CSV (if you want to keep individual converted files), and collect the data into a list for merging:
# Initialize a list to store all individual DataFrames all_customer_data = [] for file in tsv_files: # Read the TSV file exactly like your single-file method df = pd.read_csv(file, sep='\t', low_memory=False) # Optional: Save each TSV as a separate CSV (remove if you don't need this) csv_file_name = os.path.splitext(file)[0] + ".csv" df.to_csv(csv_file_name, index=False) # Add the current file's data to our list all_customer_data.append(df)
Step 4: Merge & Save the Final Large CSV
Combine all the collected DataFrames into one, then save it as a single consolidated CSV:
# Merge all DataFrames into one (ignore_index resets row numbers) merged_data = pd.concat(all_customer_data, ignore_index=True) # Save the merged dataset to your folder merged_data.to_csv(os.path.join(folder_path, "Merged_CustomerData.csv"), index=False)
Pro Tips for Smooth Processing
- Memory Management: If your files are huge, specify
dtypefor columns (e.g.,dtype={"customer_id": str}) when reading to cut down on memory usage. For extra large files, usechunksizeinpd.read_csvto process data in smaller batches. - Column Consistency: Make sure all TSV files have the same column structure—if not,
pd.concatwill add missing columns filled withNaN. Add a quick check in the loop (likeprint(df.columns)) if you suspect mismatches. - Cross-Platform Compatibility: Using
os.path.joinensures your code works on Windows, macOS, and Linux without path separator issues.
内容的提问来源于stack exchange,提问作者f722axzo5d

