You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中访问并增强图像(如锐化)且保留原结构?

如何在保留DataFrame结构的前提下处理图像(如锐化)

你可以通过两种方式实现需求:一种是手动读取并修改图像文件,同时保留原DataFrame结构;另一种是结合Keras的图像生成器实时处理图像,无需修改本地文件。以下是具体实现:

方案一:手动读取+处理图像,保留原DataFrame结构

这种方式会直接修改本地图像文件(或保存到新路径),DataFrame的结构完全不变,路径列始终对应处理后的图像。

步骤:

  1. 导入所需库
  2. 定义图像锐化函数
  3. 遍历DataFrame中的图像路径,处理图像并保存
import pandas as pd
import cv2
import numpy as np

# 定义锐化函数
def sharpen_image(image):
    # 构建锐化卷积核
    kernel = np.array([[0, -1, 0],
                       [-1, 5,-1],
                       [0, -1, 0]])
    # 应用锐化效果
    sharpened_img = cv2.filter2D(image, -1, kernel)
    return sharpened_img

# 假设你已经通过现有代码生成了DataFrame df
# 遍历每一行处理图像
for idx, row in df.iterrows():
    img_path = row['Image_name']
    # 读取图像
    img = cv2.imread(img_path)
    if img is not None:
        # 执行锐化
        processed_img = sharpen_image(img)
        # 选择1:覆盖原图像
        cv2.imwrite(img_path, processed_img)
        
        # 选择2:保存到新路径并更新DataFrame(可选)
        # new_img_path = img_path.replace('train', 'train_sharpened')
        # cv2.imwrite(new_img_path, processed_img)
        # df.at[idx, 'Image_name'] = new_img_path

处理完成后,原DataFrame的列结构、标签列完全保留,仅图像文件被更新(或路径指向新的处理后图像)。

方案二:结合ImageDataGenerator自定义预处理(实时处理)

如果不需要永久修改本地图像,可以在生成训练批量数据时实时应用锐化处理,DataFrame结构全程不变。

修改你现有的代码,添加自定义预处理函数:

from tensorflow.keras.preprocessing.image import ImageDataGenerator
import pandas as pd
import cv2
import numpy as np

# 定义锐化预处理函数
def sharpen_image(image):
    # ImageDataGenerator读取的图像是RGB格式,需转换为cv2常用的BGR格式
    image = cv2.cvtColor(image, cv2.COLOR_RGB2BGR)
    # 锐化核
    kernel = np.array([[0, -1, 0],
                       [-1, 5,-1],
                       [0, -1, 0]])
    sharpened_img = cv2.filter2D(image, -1, kernel)
    # 转换回RGB格式供模型使用
    return cv2.cvtColor(sharpened_img, cv2.COLOR_BGR2RGB)

def load_train_generator(imgs_info='/content/drive/My Drive/phd/new_data/train.xlsx',imgs_dir='/content/drive/My Drive/phd/new_data/train',
                         x_col="Image_name", y_cols="Plane", shuffle=False, batch_size=32, seed=1, target_w=256, target_h=256):

    df = pd.read_excel(imgs_info, index_col=0, engine='openpyxl')
    df["Image_name"] = df.index
    df.reset_index(drop=True, inplace=True)
    df.Image_name = imgs_dir + "/" + df.Image_name
    df = df.sample(frac=1).reset_index(drop=True)

    # 初始化图像生成器,加入自定义预处理
    image_generator = ImageDataGenerator(
        rotation_range=40,
        samplewise_center=False,
        samplewise_std_normalization=False,
        preprocessing_function=sharpen_image  # 绑定锐化函数
    )

    train_generator=image_generator.flow_from_dataframe(
        dataframe=df, 
        directory=None, 
        x_col=x_col, 
        y_col=y_cols, 
        subset=None,  # 未设置validation_split时无需指定subset
        batch_size=batch_size, 
        seed=seed, 
        shuffle=False, 
        class_mode="categorical", 
        target_size=(target_w, target_h)
    )

    return train_generator, df

注意:

  • 原代码中的subset="training"需要删除,因为你没有在ImageDataGenerator中设置validation_split,否则会触发报错。
  • 这种方式不会修改本地图像文件,仅在生成训练数据时实时处理,适合数据增强场景。

内容的提问来源于stack exchange,提问作者Amene Vatanparast

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 08:50:41