如何在Python中基于列表为浮游生物CSV文件新增列?
Got it, let's work through this problem step by step. You've got a CSV (VV_AL_3T3_P3.csv) linking particle data to TIFF images, and those images have since been sorted into shape-based folders. We need to add a new column to the CSV that tells us which folder each particle's image lives in.
Prerequisites
First, make sure you have Python's pandas library installed—it's perfect for this kind of table manipulation:
pip install pandas
Option 1: Basic Folder Search (Great for Small Datasets)
This approach searches through your categorized folders for each image filename and returns its parent folder name.
import pandas as pd import os def get_image_folder(image_name, root_dir): # Traverse all subfolders under the root directory for dirpath, _, filenames in os.walk(root_dir): if image_name in filenames: # Return just the folder name (use dirpath if you need the full path) return os.path.basename(dirpath) # If the image isn't found, return a fallback value return "Unknown" # Load your original CSV df = pd.read_csv("VV_AL_3T3_P3.csv") # Replace this with the actual root folder holding your categorized image folders image_root = "/your/path/to/categorized/image/folders" # Add the new column to the DataFrame df["Image_Folder"] = df["Image_File"].apply(lambda img: get_image_folder(img, image_root)) # Save the updated CSV (we use a new filename to avoid overwriting the original) df.to_csv("VV_AL_3T3_P3_updated.csv", index=False)
Option 2: Optimized Mapping (Better for Large Datasets)
If you have hundreds or thousands of images, the above method can be slow because it searches folders for every single row. Instead, we'll build a lookup dictionary first (only traverse folders once):
import pandas as pd import os # Build a dictionary mapping each image filename to its folder image_folder_map = {} image_root = "/your/path/to/categorized/image/folders" for dirpath, _, filenames in os.walk(image_root): current_folder = os.path.basename(dirpath) for filename in filenames: image_folder_map[filename] = current_folder # Load and update the CSV df = pd.read_csv("VV_AL_3T3_P3.csv") df["Image_Folder"] = df["Image_File"].map(image_folder_map).fillna("Unknown") # Save the result df.to_csv("VV_AL_3T3_P3_updated.csv", index=False)
Key Notes
- Filename Matching: Double-check that the
Image_Filevalues in your CSV exactly match the filenames in your folders (case-sensitive on Linux/macOS, not on Windows). - Duplicate Images: If an image exists in multiple folders, the first method returns the first folder found; the second method will use the last folder encountered during the walk. Adjust the logic if you need to handle duplicates differently.
- Fallback Value: We use "Unknown" for images that can't be found—you can change this to whatever makes sense for your workflow.
内容的提问来源于stack exchange,提问作者Olga

