如何基于条件在Pandas DataFrame两列间选择值生成新列
Pandas 根据条件生成新列问题
给定如下DataFrame:
import pandas as pd d = {'one': pd.Series([10, 20, 30, 40], index=['a', 'b', 'c', 'd']), 'two': pd.Series([0, 0, 25, 0], index=['a', 'b', 'c', 'd'])} df = pd.DataFrame(d)
生成的DataFrame内容为:
one two a 10 0 b 20 0 c 30 25 d 40 0
需要根据以下逻辑生成新列df['three']:当df['two'] > 0时取df['two']的值,否则取df['one']的值,最终期望结果如下:
one two three a 10 0 10 b 20 0 20 c 30 25 25 d 40 0 40
解决方案
方法1:使用numpy.where(推荐,向量式操作效率高)
适合处理大规模数据集,直接对整列进行条件判断和赋值:
import numpy as np df['three'] = np.where(df['two'] > 0, df['two'], df['one'])
方法2:使用Pandas原生where方法
Pandas的where会保留满足条件的原值,不满足条件的替换为指定值,刚好适配需求逻辑:
df['three'] = df['two'].where(df['two'] > 0, df['one'])
方法3:使用apply逐行处理(不推荐大规模数据)
如果需要逐行逻辑处理可以用,但效率远低于前两种向量式方法:
df['three'] = df.apply(lambda row: row['two'] if row['two'] > 0 else row['one'], axis=1)
内容的提问来源于stack exchange,提问作者enter_display_name_here
相关产品推荐
相关产品推荐

