You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中基于两列多条件创建新列并简化冗余代码

优化Pandas条件判断代码:简化多值匹配逻辑

问题背景

需要基于数据集的condition_code和trans_type两列组合值,生成新列标记会计交易为"posted"或"not posted"。现有代码通过大量.eq()和|操作符拼接条件,代码冗长近200行。尝试用Python原生in操作符匹配列表值时,因返回整列布尔值而非逐行判断失败,需寻求简洁可行的优化方案。

现有冗长代码示例

conditions = [df['trans_type'].eq('D67'), 
  (df['condition_code'].eq('H')) & 
  (df['trans_type'].eq('D4S') | df['trans_type'].eq('D4U') | ...  # 大量.eq()拼接
  many more .eq statements),
  other conditions in the same format...]

错误尝试代码

conditions = [df['trans_type'] == 'D67', 
  (df['condition_code'] in ['A', 'H']) & (df['trans_type'] in 
  ['D4S', 'D4U', 'D4V', ...]),  # in操作符不适用于Pandas列的逐行判断
  more similar conditions...]

可行优化方案

方案1:使用Pandas .isin() 实现向量化多值匹配

Pandas的.isin()方法支持逐行判断列值是否在目标列表中,完美替代多个.eq()+|的组合,大幅简化代码:

import numpy as np

# 用isin替代多个eq+|,重构条件列表
conditions = [
    df['trans_type'].isin(['D67']),
    (df['condition_code'].isin(['A', 'H'])) & (df['trans_type'].isin(['D4S', 'D4U', 'D4V'])),
    # 其他条件同理替换:把所有同类型的eq+|改成isin
]

# 定义每个条件对应的结果,最后用np.select生成新列
choices = ['posted', 'posted', ...]  # 每个condition对应结果,顺序与conditions一致
df['post_status'] = np.select(conditions, choices, default='not posted')

优势:向量化操作,性能远高于循环/逐行判断,适合处理大型数据集。

方案2:用集合存储合法组合,结合apply逐行判断

如果判断逻辑是基于condition_code和trans_type的固定组合,可以将所有需要标记为"posted"的组合存入集合,再通过apply逐行校验:

# 定义所有应标记为posted的(condition_code, trans_type)组合
posted_pairs = {
    ('', 'D67'),  # 对应仅trans_type为D67的情况
    ('H', 'D4S'),
    ('H', 'D4U'),
    ('A', 'D4V'),
    # 所有其他合法组合
}

# 逐行判断当前行的组合是否在集合中
df['post_status'] = df.apply(
    lambda row: 'posted' if (row['condition_code'], row['trans_type']) in posted_pairs else 'not posted',
    axis=1
)

优势:代码逻辑直观,组合规则维护方便;缺点是apply属于逐行操作,大数据集下性能不如方案1。

内容的提问来源于stack exchange,提问作者X8bitReignbeaux

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 23:17:07