You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于行值计算创建DataFrame新列的高效方法

Pandas高效生成计算列(替代iterrows)

先看原始数据和需求:
我们需要基于每行c2到c6的切片均值生成三个新列,避开低效的iterrows方法。

原始代码:

import pandas as pd
import numpy as np

data = {"c1": [10], "c2": [20], "c3":[30], "c4":[40], "c5":[50], "c6":[10]}
df = pd.DataFrame(data=data)

方法1:用apply(比iterrows高效)

如果逻辑相对复杂,可使用apply逐行处理,效率优于iterrows:

def calc_row(row):
    s = row[['c2','c3','c4','c5','c6']].values
    col1 = np.mean(s[:2])
    col2 = np.mean(s[2:4])
    col3 = col1 + col2
    return pd.Series([col1, col2, col3], index=['new_column1','new_column2','new_column3'])

# 把计算结果合并到原表
df[['new_column1','new_column2','new_column3']] = df.apply(calc_row, axis=1)

方法2:向量化操作(效率最高,优先推荐)

如果逻辑可拆分为整列操作,直接用向量化处理——这是Pandas中性能最优的方式,完全无需逐行遍历:

# 直接对指定列求行均值
df['new_column1'] = df[['c2','c3']].mean(axis=1)
df['new_column2'] = df[['c4','c5']].mean(axis=1)
# 新列3直接用前两个计算列相加
df['new_column3'] = df['new_column1'] + df['new_column2']

最终输出结果:

c1  c2  c3  c4  c5  c6  new_column1  new_column2  new_column3
0  10  20  30  40  50  10         25.0         45.0         70.0

内容的提问来源于stack exchange,提问作者tudou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 17:31:00