批量将成对CSV文件合并为单个文件的简便方法
Automate Merging Paired CSV Files
Great question! Instead of manually updating filenames each time, we can use Python's os module to scan your directory, group files by their base name (like aapl, aa), and automatically merge each pair. Here's a robust solution:
Step-by-Step Solution
1. Full Code Implementation
import pandas as pd import os # Replace this with the path to your directory containing the CSV files directory = "/path/to/your/csv/files" # Create a dictionary to group files by their base name (e.g., 'aapl', 'aa') file_groups = {} # Iterate over all files in the directory for filename in os.listdir(directory): if filename.endswith(".csv"): # Split the filename to extract the base name filename_parts = filename.split("-") if len(filename_parts) >= 3: base_name = filename_parts[0] # Add the file to its corresponding group if base_name not in file_groups: file_groups[base_name] = [] file_groups[base_name].append(filename) # Process each group of files for base, files in file_groups.items(): bal_file = None cas_file = None # Find the BAL and CAS files in the group for file in files: if "BAL-Q.csv" in file: bal_file = file elif "CAS-Q.csv" in file: cas_file = file # Only proceed if both files exist if bal_file and cas_file: # Read both CSV files df_bal = pd.read_csv(os.path.join(directory, bal_file)) df_cas = pd.read_csv(os.path.join(directory, cas_file)) # Merge the DataFrames (same logic as your original code) merged_df = pd.concat( [df_bal, df_cas], join="outer", axis=0, ignore_index=True ) # Save the merged file output_filename = f"{base}-ALL.csv" merged_df.to_csv(os.path.join(directory, output_filename), index=False) print(f"Successfully merged: {bal_file} + {cas_file} → {output_filename}") else: # Handle cases where one file is missing missing_files = [] if not bal_file: missing_files.append(f"{base}-BAL-Q.csv") if not cas_file: missing_files.append(f"{base}-CAS-Q.csv") print(f"Skipping {base}: Missing files - {', '.join(missing_files)}")
2. Key Features Explained
- Automatic File Grouping: The script scans your directory and groups files by their base name (e.g., all files starting with
aaplare grouped together). - Error Handling: It checks if both
BAL-QandCAS-Qfiles exist for each base, so you won't get errors from missing files. - Reusable: Just update the
directoryvariable to point to your files, and it will process all 200 pairs in one go. - Consistent Logic: Uses the same
pd.concatparameters you originally used (withjoin='outer'andignore_index=True), so the merged output matches your manual results.
3. Notes for Usage
- Make sure you have pandas installed (
pip install pandasif not). - Replace
/path/to/your/csv/fileswith the actual path to your directory (e.g.,C:/Users/You/Documents/CSVFileson Windows, or/home/you/csv_fileson Linux/macOS). - The script will save merged files like
aapl-ALL.csvdirectly in the same directory as the source files.
内容的提问来源于stack exchange,提问作者dejanmarich
相关产品推荐
相关产品推荐

