You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何循环遍历Pandas DataFrame列并为每列创建新DataFrame执行函数

遍历Pandas DataFrame列并逐列处理的优化实现

你需要遍历一个多列的Pandas DataFrame,为每列生成单独的DataFrame后执行处理函数,以下是优化的实现方案:

示例数据

date                   Col1      Col2       Col3      Col4           
1990-01-02 12:00:00     24        24        24.8      24.8           
1990-01-02 01:00:00     59        58        60        60.3   
1990-01-02 02:00:00     43.7      43.9      48        49

现有单列处理代码

df_new = pd.DataFrame(df['Col1'])
df.reset_index(inplace=True)

def function1(df_new):
    line 1
    line 2

def function2():
    line 1
    line 2

优化后的循环实现

你的思路方向是对的,这里给出更稳健、易维护的实现方式:

import pandas as pd

# 假设你的原始DataFrame为df,date列为索引
processed_data = {}

# 整合处理逻辑到单个函数,方便复用和维护
def process_column(col_df, col_name):
    # 重置索引,避免inplace操作修改原数据
    col_df = col_df.reset_index()
    # 执行function1的逻辑
    # function1(col_df)
    # 执行function2的逻辑(如需参数可调整函数定义)
    # function2()
    # 示例操作:添加列名标识(可按需删除)
    col_df['processed_column'] = col_name
    return col_df

# 遍历所有列
for col in df.columns:
    # 生成仅包含当前列的DataFrame
    single_col_df = df[[col]]
    # 处理当前列
    result_df = process_column(single_col_df, col)
    # 保存处理结果(不需要可删除此步骤)
    processed_data[col] = result_df

# 查看处理后的结果示例
print(processed_data['Col1'])

关键优化点

  • 避免inplace操作:reset_index()不使用inplace=True,而是重新赋值,防止意外修改原DataFrame的数据
  • 封装处理逻辑:将分散的function1、function2整合到单个处理函数中,代码结构更清晰,便于后续修改和调试
  • 结果存储:用字典存储每列的处理结果,方便后续按列名快速调用,无需重复处理
  • 变量一致性:注意原框架中df_full应为df(假设是笔误),确保循环中变量名统一

内容的提问来源于stack exchange,提问作者A Newbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 04:40:55