如何读取文件夹中各文件并创建对应名称的独立DataFrame?
Instead of concatenating all files into a single DataFrame, you can use a dictionary to store each DataFrame under a key corresponding to the file name. This is the cleanest and most maintainable approach in Python.
Recommended Approach: Use a Dictionary
Here's how to modify your code to create separate DataFrames stored in a dictionary:
import glob import pandas as pd path = r'C:\Users\SemR\Documents\Jupyter\Submissions' all_files = glob.glob(path + "/*.csv") # Initialize an empty dictionary to hold your DataFrames df_dict = {} for filename in all_files: # Extract the base filename (without path and .csv extension) # Split on backslashes (Windows path) and take the last part, then remove .csv file_name = filename.split('\\')[-1].replace('.csv', '') # Read the CSV into a DataFrame, keeping only your desired columns df = pd.read_csv(filename, index_col=None, header=0, usecols=['Date', 'Usage']) # Store the DataFrame in the dictionary using the filename as the key df_dict[file_name] = df
How to Access Your DataFrames
Once the code runs, you can access each DataFrame by its filename using the dictionary:
# Example: Access the DataFrame for a file named "January_Data.csv" df_dict['January_Data']
You can list all available DataFrame names with:
print(df_dict.keys())
Alternative: Create Individual Variables (Not Recommended)
If you specifically want separate variables named after each file (rather than using a dictionary), you can use globals() to dynamically create variables. However, this approach is not ideal because it can clutter your namespace, cause name conflicts, and make it harder to manage multiple DataFrames.
Here's how you'd do it:
import glob import pandas as pd path = r'C:\Users\SemR\Documents\Jupyter\Submissions' all_files = glob.glob(path + "/*.csv") for filename in all_files: file_name = filename.split('\\')[-1].replace('.csv', '') df = pd.read_csv(filename, index_col=None, header=0, usecols=['Date', 'Usage']) # Dynamically create a variable with the filename globals()[file_name] = df
Warnings for This Approach:
- If your filenames contain spaces, hyphens, or special characters, they'll become invalid variable names (e.g., "January Data.csv" would try to create a variable named
January Data, which is not allowed). - You risk overwriting existing variables if a filename matches an existing variable name.
- It's difficult to track all the DataFrames you've created without manually checking the namespace.
Stick with the dictionary approach for a cleaner, more manageable solution.
内容的提问来源于stack exchange,提问作者R Sem

