如何在Pandas读取多文件时添加文件名作为单独列?
Solution: Add Filename Column When Combining Multiple CSVs with Pandas
Hey there! Let's get that filename column added to your combined DataFrame—you're almost there, just a couple small tweaks needed.
First, let's fix a tiny typo in your existing code: you wrote irnore_index = TRUE but it should be ignore_index=True (Python uses lowercase boolean values, and the spelling was off).
Now, to add the filename column, you'll want to:
- Import the
osmodule to easily extract filenames from full file paths - For each file you read, add a new column to the DataFrame that stores the filename (or full path, if you prefer)
Here's the updated, working code:
import pandas as pd import glob import os # Add this module to parse file paths path = r'C:\my_file_path\Misc' allFiles = glob.glob(path + "/*.csv") list_ = [] for file_ in allFiles: df = pd.read_csv(file_, index_col=None, dtype=str, header=0) # Add a column with just the filename (e.g., "sales_data.csv") df['source_filename'] = os.path.basename(file_) # If you want the full file path instead, use this line: # df['source_filepath'] = file_ list_.append(df) # Fixed the typo here for ignore_index frame = pd.concat(list_, axis=0, ignore_index=True)
Quick breakdown:
os.path.basename(file_)takes a full path likeC:\my_file_path\Misc\Q3_sales.csvand pulls out just the filename part (Q3_sales.csv)—this keeps your data clean while tracking where each row came from.- Adding this column before appending the DataFrame to your list ensures every row in every file gets tagged with its source. When you concatenate all DataFrames together, this column will carry over to your final combined dataset.
That's all you need! Your frame DataFrame will now have all your CSV data plus a clear marker for which file each row originated from.
内容的提问来源于stack exchange,提问作者Bryan_UK
相关产品推荐
相关产品推荐

