You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pd.DataFrame.apply传递参数时遇结果异常问题

问题:DataFrame.apply调用增强函数后目标DF出现错误

我想给原DataFrame的每一行生成增强数据并写入新DataFrame,为此定义了augment函数。单独调用该函数一切正常,但用pd.DataFrame.apply传参调用后,目标DataFrame出现<Error>错误,请问问题出在哪?

函数代码

def augment(row: pd.Series, column_name: str, target_df: pd.DataFrame, num_samples: int):
    # print(type(row))
    target_df_start_index = target_df.shape[0]
    raw_img = row[column_name].astype('uint8')
    bin_image = convert_image_to_binary_image(raw_img)
    bin_3dimg = tf.expand_dims(input=bin_image, axis=2)
    bin_img_reshaped = tf.image.resize_with_pad(image=bin_3dimg, target_width=128, target_height=128, method="bilinear")

    for i in range(num_samples + 1):
        new_row = row.copy(deep=True)

        if i == 0:
            new_row[column_name] = np.squeeze(bin_img_reshaped, axis=2)
        else:
            aug_image = data_augmentation0(bin_img_reshaped)
            new_row[column_name] = np.squeeze(aug_image, axis=2)

        # display.display(new_row)
        target_df.loc[target_df_start_index + i] = new_row

    # print(target_df.shape)
    # display.display(target_df)

单独调用示例

tmp_df = pd.DataFrame(None, columns=testDF.columns)
augment(testDF.iloc[0], column_name='binMap', target_df=tmp_df, num_samples=4)
augment(testDF.iloc[1], column_name='binMap', target_df=tmp_df, num_samples=4)

apply调用示例

tmp_df = pd.DataFrame(None, columns=testDF.columns)
testDF.apply(augment, args=('binMap', tmp_df, 4, ), axis=1)

调用后错误结果

,data
,
,


问题原因及解决方案

核心原因

  1. apply的设计逻辑不匹配:apply的初衷是处理每行数据并返回结果,而非原地修改外部DataFrame。你这种在函数里直接操作外部tmp_df的方式属于副作用操作,容易触发Pandas内部的状态异常。
  2. 返回值干扰:augment函数没有明确返回值,默认返回None。apply会自动收集每行的返回值生成新对象,这个过程可能破坏tmp_df的内部结构,导致显示<Error>。
  3. 索引与数据结构冲突:apply处理行时会对Series做内部包装,可能导致new_row的结构和tmp_df的列定义不兼容,赋值后引发错误。

解决方案

方案1:改用普通循环(最稳妥)

既然单独调用函数正常,直接遍历原DataFrame的行调用augment即可,逻辑和单独调用完全一致:

tmp_df = pd.DataFrame(None, columns=testDF.columns)
for idx, row in testDF.iterrows():
    augment(row, column_name='binMap', target_df=tmp_df, num_samples=4)

方案2:让函数返回增强行,再合并到目标DF

修改augment函数,让它返回当前行生成的所有增强行,再通过apply收集结果并拼接成tmp_df,完全符合apply的设计逻辑:

def augment(row: pd.Series, column_name: str, num_samples: int):
    raw_img = row[column_name].astype('uint8')
    bin_image = convert_image_to_binary_image(raw_img)
    bin_3dimg = tf.expand_dims(input=bin_image, axis=2)
    bin_img_reshaped = tf.image.resize_with_pad(image=bin_3dimg, target_width=128, target_height=128, method="bilinear")
    
    augmented_rows = []
    for i in range(num_samples + 1):
        new_row = row.copy(deep=True)
        if i == 0:
            new_row[column_name] = np.squeeze(bin_img_reshaped, axis=2)
        else:
            aug_image = data_augmentation0(bin_img_reshaped)
            new_row[column_name] = np.squeeze(aug_image, axis=2)
        augmented_rows.append(new_row)
    
    return augmented_rows

# 调用方式
augmented_list = testDF.apply(augment, args=('binMap', 4, ), axis=1).explode().tolist()
tmp_df = pd.DataFrame(augmented_list, columns=testDF.columns)

方案3:强制忽略apply的返回结果(不推荐)

如果一定要用apply,可以在函数末尾返回None,并主动丢弃apply的返回值,减少副作用影响,但这种方式仍存在不可预测的风险:

def augment(row: pd.Series, column_name: str, target_df: pd.DataFrame, num_samples: int):
    # 原有代码不变...
    
    # 明确返回None
    return None

# 调用时忽略apply的返回结果
tmp_df = pd.DataFrame(None, columns=testDF.columns)
_ = testDF.apply(augment, args=('binMap', tmp_df, 4, ), axis=1)

内容的提问来源于stack exchange,提问作者soumeng78

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 01:09:28