如何用Pandas函数根据列非空值生成指定支付类型列
问题分析与解决指引
你的代码存在三个核心问题:
- 多余的循环逻辑:使用
apply(axis=1)时,传入函数的x是单行数据的Series,x['Loan Funded Date'].notna()返回的是单个布尔值,而非可迭代的序列,因此嵌套for循环完全没必要,反而会导致逻辑混乱。 - 条件优先级错误:当前逻辑中如果两列都非空,
Insurance会覆盖Financing,不符合你先判断贷款日期的需求。 - 缺少两列均为空时的赋值逻辑:代码里没有处理最后一种情况的分支。
修正方案一:优化原apply函数
直接针对单行数据做条件判断,去掉多余循环:
import pandas as pd def typepayment(x): if pd.notna(x['Loan Funded Date']): return 'Financing' elif pd.notna(x['Claim Approved Date']): return 'Insurance' else: return 'Cash/Credit' # 应用函数生成新列 df['Payment Type'] = df.apply(typepayment, axis=1)
修正方案二:更高效的矢量化操作(推荐)
Pandas中apply在大数据量下效率较低,推荐使用矢量化方法,比如numpy.select或Pandas的loc赋值:
方法1:使用numpy.select
import numpy as np # 定义条件列表和对应结果 conditions = [ df['Loan Funded Date'].notna(), df['Claim Approved Date'].notna() ] choices = [ 'Financing', 'Insurance' ] # 生成新列,默认值为Cash/Credit df['Payment Type'] = np.select(conditions, choices, default='Cash/Credit')
方法2:使用Pandas loc链式赋值
# 先设置默认值 df['Payment Type'] = 'Cash/Credit' # 覆盖保险审批日期非空的情况 df.loc[df['Claim Approved Date'].notna(), 'Payment Type'] = 'Insurance' # 最后覆盖贷款放款日期非空的情况(保证优先级) df.loc[df['Loan Funded Date'].notna(), 'Payment Type'] = 'Financing'
内容的提问来源于stack exchange,提问作者internet joe
相关产品推荐
相关产品推荐

