在Google Data Prep导入Google Storage文件夹时,如何将文件名及元数据设为数据集列
Hey there! I’ve done this exact workflow in Google Data Prep a bunch of times, so let me walk you through how to make it happen smoothly. Here’s your step-by-step guide:
Step 1: Import Your Entire GCS Folder
First up, when creating a new dataset in Data Prep:
- Pick Google Cloud Storage as your source connection.
- Navigate to the folder you want to import (don’t click on individual files!) and select it. Data Prep will automatically pull in all compatible files in that folder—no need to select them one by one.
Step 2: Add the Source Filename as a Column
This is straightforward once your data loads:
- In the top toolbar, hit the Add Column dropdown.
- Hover over Add from Metadata, then select Filename. A new column (usually named
filename) will pop up, populated with either the full GCS path or just the base name of the file each row originated from. - If you only want the base filename (without the full path), use a quick Split transform: split the
filenamecolumn on/and keep the last segment. Or use a custom formula likesplit(filename, '/')[ -1 ]to strip the path in one go.
Step 3: Bring in Additional File Metadata Columns
Want more than just the filename? You can add other GCS file metadata as columns too:
- Go back to Add Column > Add from Metadata, and pick any of the available fields. Common useful options include:
file_size: The size of the source file (in bytes)last_modified: Timestamp when the file was last updated in GCSfile_path: Full GCS path to the source file
- Each selection adds a new column with that metadata—super intuitive.
Quick Pro Tips to Avoid Headaches
- If your folder has files with different schemas, adjust the Schema Inference settings first (you’ll see this during import) to merge schemas or handle mismatches. Otherwise, metadata columns might not align correctly.
- If your GCS folder has nested subfolders, check the Include subfolders box during import if you want those files included too.
- Feel free to rename metadata columns (like changing
last_modifiedtofile_updated_at) using the Rename transform—it makes your final dataset much cleaner for downstream use.
内容的提问来源于stack exchange,提问作者B-Rad
相关产品推荐
相关产品推荐

