You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于两列条件为Pandas DataFrame生成新列col6

解决Pandas按条件生成新列的问题

嗨,我来帮你搞定这个生成新列的需求!根据你给出的规则,我们可以用几种高效的方式实现,下面给你详细说明:


方法1:使用numpy.where()(推荐,矢量化操作效率高)

这是最适合的方法,因为矢量化操作在Pandas中处理大数据集时速度最快。我们可以嵌套where来处理三个分支的条件:

import pandas as pd
import numpy as np

# 构造你的示例DataFrame
df = pd.DataFrame({
    'col1': ['abc', 'abc', 'def', 'def'],
    'col2': [1, 2, 1, 2],
    'col3': [20, 25, 340, 185],
    'col4': [10, 5, 12, 16],
    'col5': [15, 30, 22, 120]
})

# 生成col6
df['col6'] = np.where(
    df['col2'] == 1,
    df['col3'] * df['col5'],
    np.where(
        df['col2'] == 2,
        df['col4'] * df['col5'],
        0
    )
)

print(df)

代码解释:

  • 第一层where判断col2 == 1,满足则计算col3 * col5
  • 不满足的话进入第二层where,判断col2 == 2,满足则计算col4 * col5
  • 两个条件都不满足时,返回0,完美兼容col2有其他值的情况

方法2:使用Pandas的loc索引(直观易懂)

如果你更喜欢分步操作,用loc给不同条件的行赋值也很清晰:

# 先初始化col6为0
df['col6'] = 0

# 给col2==1的行赋值
df.loc[df['col2'] == 1, 'col6'] = df['col3'] * df['col5']

# 给col2==2的行赋值
df.loc[df['col2'] == 2, 'col6'] = df['col4'] * df['col5']

这种方法逻辑直白,适合新手理解,同样是矢量化操作,效率也很高。


方法3:使用apply()(适合小数据集)

如果你的数据集不大,也可以用apply逐行处理,但要注意大数据量下效率不如前两种方法:

def calculate_col6(row):
    if row['col2'] == 1:
        return row['col3'] * row['col5']
    elif row['col2'] == 2:
        return row['col4'] * row['col5']
    else:
        return 0

df['col6'] = df.apply(calculate_col6, axis=1)

最终输出结果

不管用哪种方法,你都会得到想要的DataFrame:

col1col2col3col4col5col6
abc1201015300
abc225530150
def134012227480
def2185161201920

内容的提问来源于stack exchange,提问作者jceg316

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:13:40