You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

数据转换优化需求:按规则生成check列(求高效实现方案)

问题描述

现有数据集需新增check列,规则如下:

  • 若consumer的master_consumer为null且bill_group为'Consumer',check值等于Consumer_scope;
  • 若master_consumer非空且存在于consumer列中,当其Consumer_scope为'out of scope'时,该master_consumer下属所有consumer的check均为'out of scope';
  • 若某consumer的product含'For Dairy'且属于某master_consumer,该master_consumer下属所有consumer的check均为'Out of Scope'。

当前实现代码性能极差,求高效实现方案。

高效实现方案(基于Pandas)

核心思路

放弃逐行循环,用向量化操作+集合查询完成批量处理,这是性能提升的核心。

具体步骤

  1. 预筛选需标记的异常主消费方

    • 提取规则2的异常主消费方:筛选consumer列中Consumer_scope为'out of scope'的行,取出这些行的consumer值存入集合:
      bad_masters_scope = set(df[df['Consumer_scope'] == 'out of scope']['consumer'])
      
    • 提取规则3的异常主消费方:筛选product包含'For Dairy'且master_consumer非空的行,取出master_consumer值存入集合:
      bad_masters_dairy = set(df[df['product'].str.contains('For Dairy') & df['master_consumer'].notnull()]['master_consumer'])
      
    • 合并两个集合,得到所有需要标记下属的主消费方:
      all_bad_masters = bad_masters_scope.union(bad_masters_dairy)
      
  2. 批量赋值check列

    • 初始化check列,先处理规则1的场景:
      df['check'] = None
      # 规则1:master为空且bill_group是Consumer的行,直接取Consumer_scope
      rule1_mask = (df['master_consumer'].isnull()) & (df['bill_group'] == 'Consumer')
      df.loc[rule1_mask, 'check'] = df.loc[rule1_mask, 'Consumer_scope']
      
    • 处理规则2和3的场景:所有属于异常主消费方的下属,统一标记为'out of scope'
      bad_master_mask = df['master_consumer'].isin(all_bad_masters)
      df.loc[bad_master_mask, 'check'] = 'out of scope'
      
    • (可选)为未覆盖的行设置默认值(根据业务需求调整):
      df['check'] = df['check'].fillna('in scope')
      

性能优化补充

  • 若数据集规模极大,可先对master_consumer列做排序+分组,减少重复查询;或者用Dask框架进行并行处理,进一步提升速度。
  • 避免使用apply()、iterrows()这类逐行处理的方法,它们的时间复杂度是O(n),远慢于向量化操作的O(1)级别。

内容的提问来源于stack exchange,提问作者Steve Timb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 17:32:48