You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中基于列表为浮游生物CSV文件新增列?

Add a Categorized Folder Column to Your Plankton CSV

Got it, let's work through this problem step by step. You've got a CSV (VV_AL_3T3_P3.csv) linking particle data to TIFF images, and those images have since been sorted into shape-based folders. We need to add a new column to the CSV that tells us which folder each particle's image lives in.

Prerequisites

First, make sure you have Python's pandas library installed—it's perfect for this kind of table manipulation:

pip install pandas

Option 1: Basic Folder Search (Great for Small Datasets)

This approach searches through your categorized folders for each image filename and returns its parent folder name.

import pandas as pd
import os

def get_image_folder(image_name, root_dir):
    # Traverse all subfolders under the root directory
    for dirpath, _, filenames in os.walk(root_dir):
        if image_name in filenames:
            # Return just the folder name (use dirpath if you need the full path)
            return os.path.basename(dirpath)
    # If the image isn't found, return a fallback value
    return "Unknown"

# Load your original CSV
df = pd.read_csv("VV_AL_3T3_P3.csv")

# Replace this with the actual root folder holding your categorized image folders
image_root = "/your/path/to/categorized/image/folders"

# Add the new column to the DataFrame
df["Image_Folder"] = df["Image_File"].apply(lambda img: get_image_folder(img, image_root))

# Save the updated CSV (we use a new filename to avoid overwriting the original)
df.to_csv("VV_AL_3T3_P3_updated.csv", index=False)

Option 2: Optimized Mapping (Better for Large Datasets)

If you have hundreds or thousands of images, the above method can be slow because it searches folders for every single row. Instead, we'll build a lookup dictionary first (only traverse folders once):

import pandas as pd
import os

# Build a dictionary mapping each image filename to its folder
image_folder_map = {}
image_root = "/your/path/to/categorized/image/folders"

for dirpath, _, filenames in os.walk(image_root):
    current_folder = os.path.basename(dirpath)
    for filename in filenames:
        image_folder_map[filename] = current_folder

# Load and update the CSV
df = pd.read_csv("VV_AL_3T3_P3.csv")
df["Image_Folder"] = df["Image_File"].map(image_folder_map).fillna("Unknown")

# Save the result
df.to_csv("VV_AL_3T3_P3_updated.csv", index=False)

Key Notes

  • Filename Matching: Double-check that the Image_File values in your CSV exactly match the filenames in your folders (case-sensitive on Linux/macOS, not on Windows).
  • Duplicate Images: If an image exists in multiple folders, the first method returns the first folder found; the second method will use the last folder encountered during the walk. Adjust the logic if you need to handle duplicates differently.
  • Fallback Value: We use "Unknown" for images that can't be found—you can change this to whatever makes sense for your workflow.

内容的提问来源于stack exchange,提问作者Olga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:26:36