如何在Pandas中优化读取两个不同文件夹内的同名文件?
Hey there! I totally get where you're coming from—writing repetitive pd.read_csv() calls for each file is tedious and hard to maintain. Let's implement a clean, scalable solution to batch process all your same-named files in one go.
Step-by-Step Solution
Here's how you can automate this workflow:
1. Import Required Libraries
First, we'll use pandas for reading the files and os for handling file paths and directory listings:
import pandas as pd import os
2. Define Folder Paths
Set up your two target folder paths (using raw strings to avoid escape character issues):
folder_2020 = r'C:\Users\2020' folder_2021 = r'C:\Users\2021'
3. Identify Common Filenames
We'll get the list of files in each folder, then find the intersection (files that exist in both folders). We'll also add a filter to only target your team_*.txt files:
# Get all filenames in each folder files_2020 = set(os.listdir(folder_2020)) files_2021 = set(os.listdir(folder_2021)) # Filter to keep only matching team files present in both folders common_files = { filename for filename in files_2020 & files_2021 if filename.startswith('team_') and filename.endswith('.txt') }
4. Batch Read Files and Store DataFrames
Use a dictionary to store pairs of DataFrames (one from each folder) mapped to their filename. This keeps your data organized and easy to access later:
# Initialize a dictionary to hold our DataFrame pairs file_dfs = {} for filename in common_files: # Construct full paths for the file in both folders path_2020 = os.path.join(folder_2020, filename) path_2021 = os.path.join(folder_2021, filename) # Read the tab-separated files into DataFrames df_2020 = pd.read_csv(path_2020, sep='\t') df_2021 = pd.read_csv(path_2021, sep='\t') # Store the pair in the dictionary file_dfs[filename] = { '2020': df_2020, '2021': df_2021 }
5. Access Your DataFrames
Now you can easily retrieve the DataFrames for any file without repeating code:
# Get the 2020 version of team_ABC.txt df_2020_abc = file_dfs['team_ABC.txt']['2020'] # Get the 2021 version of team_XYZ.txt df_2021_xyz = file_dfs['team_XYZ.txt']['2021']
Why This Works
- Scalability: Add new
team_*.txtfiles to both folders, and the code will automatically process them without changes. - Cleaner Code: No more repetitive
pd.read_csv()lines—everything is handled in a single loop. - Organized Storage: The dictionary keeps your DataFrames grouped by filename and year, making it easy to manage and access data.
内容的提问来源于stack exchange,提问作者LearnerBegineer

