You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将逐行遍历pandas DataFrame的函数改写为apply方法提升性能

逐行遍历pandas DataFrame函数改写为apply实现方案

需求背景

现有逐行遍历pandas DataFrame运行的自定义函数,希望改写为pandas apply方法或同类实现,提升运行性能,无需手动遍历DataFrame索引。

样例数据结构

df
       object_id  ...                        param_dict
8804       15563  ...                         {81: 2.0}
8805       15566  ...                         {81: 2.0}
8806       15553  ...                         {81: 2.0}
8808       15531  ...                         {81: 2.0}
8811       15639  ...                         {81: 2.0}
...          ...  ...                               ...
16525       1158  ...  {4: 9963.302345992126, 46: 92.4}
16526       1156  ...               {4: -0.0, 46: 67.5}
16527       1089  ...                 {4: -0.0, 46: 76}
16528        898  ...               {4: -0.0, 46: 67.5}
16531        893  ...               {4: -0.0, 46: 67.5}
[1333 rows x 8 columns]

原有实现代码

def function(df):
    # running over the index of the dataframe
    for index in df.index:

        # running over the keys of the dataframe['param_dict'] dictionaries
        for key in df['param_dict'][index]:
            if df['param_dict'][index][key] == 0:
                continue

            if key in [4, 27]:
                print(df['name'][index], df['param_dict'][index][key], 1)

            elif key in [46, 28, 29]:
                print(df['name'][index], df['param_dict'][index][key], 2)

            else:
                print(df['name'][index], df['param_dict'][index][key], 3)

    return None

测试数据集

df = pd.DataFrame({'name': ['a', 'b', 'c'], 'param_dict': [{4: 0, 1: 4}, {46: True}, {35: False, 25: 0}]})

改写后的apply实现方案

步骤1:定义单行处理函数

抽取出单条数据的处理逻辑,输入为DataFrame的单行数据:

def process_single_row(row):
    current_name = row['name']
    param_dict = row['param_dict']
    for key, value in param_dict.items():
        # 跳过值为0的项
        if value == 0:
            continue
        # 按key所属分类输出
        if key in [4, 27]:
            print(current_name, value, 1)
        elif key in [46, 28, 29]:
            print(current_name, value, 2)
        else:
            print(current_name, value, 3)

步骤2:调用apply方法逐行执行

# axis=1 表示按行遍历
df.apply(process_single_row, axis=1)

运行结果验证

使用提供的测试数据集运行,输出结果与原有循环实现完全一致:

a 4 3
b True 2

性能说明

apply方法内部已经实现了遍历逻辑,无需手动处理索引,代码简洁度更高。如果需要进一步提升性能,可以考虑将param_dict列展开为结构化列后使用向量化操作,适合数据量更大的场景。

内容的提问来源于stack exchange,提问作者oakca

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 19:27:06