You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas的else分支中按H_cnt分配行至最近IF条件块

问题描述

现有如下结构的Pandas DataFrame:

customer_id   recency frequency H_cnt  Years_with_us 
   1           0         143      3       0.32    
   2           14        190      8       1.7

核心需求

当某行数据不匹配任何已定义的IF条件时,需基于H_cnt字段将该行分配到最近的IF条件对应的分组。例如:行1的H_cnt=3不匹配任何现有条件,需将其分配至H_cnt>=4对应的「Short Tenure - Promising」分组。

当前代码

cust = 'customer_id'
for row in df.iterrows():
    rec = row[1] 
    r = rec['recency'] 
    f = rec['frequency'] 
    y = rec['years_with_us'] 
    h = rec['H_cnt']
    if ((r <= 11) and (f >=131) and (y >= 0.6) and (h >= 9)):
        classes.append({rec[cust]:'Champions'})
    elif ((r <= 11) and (f >=19) and (y < 0.6) and (h >= 4)):
        classes.append({rec[cust]:'Short Tenure - Promising'})
    elif ((r <= 62) and (f >=19) and (y >= 1.5)):
        classes.append({rec[cust]:'Loyal Customers'})
    elif ((r <= 62) and (f >=19) and (y >= 0.6) and (y < 1.5)):
        classes.append({rec[cust]:'Potential Loyalist'})
    elif ((r <= 62) and (f <=18) and (y >= 0.6)):
        classes.append({rec[cust]:'New Customers'})
    else:
        print("hi")
        print(row[1])
        classes.append({0:[row[1]['recency'],row[1]['frequency'],row[1]['H_cnt'],row[1]['years_with_IFX']]})
    accs = [list(i.keys())[0] for i in classes]
    segments = [list(i.values())[0] for i in classes]
    df['new_segment'] = df[cust].map(dict(zip(accs,segments)))

期望输出

customer_id   recency frequency H_cnt  Years_with_us  new_segment
   1           0         143      3       0.32        Short Tenure - Promising
   2           14        190      8       1.7         Champions

解决方案

修改后的完整代码

cust = 'customer_id'
# 预定义带H_cnt阈值的规则及对应分组,按阈值从高到低排序
h_based_rules = [
    {'threshold': 9, 'segment': 'Champions'},
    {'threshold': 4, 'segment': 'Short Tenure - Promising'}
]

classes = []

for row in df.iterrows():
    rec = row[1] 
    r = rec['recency'] 
    f = rec['frequency'] 
    y = rec['years_with_us'] 
    h = rec['H_cnt']
    is_matched = False
    
    # 原有条件判断逻辑
    if ((r <= 11) and (f >=131) and (y >= 0.6) and (h >= 9)):
        classes.append({rec[cust]: 'Champions'})
        is_matched = True
    elif ((r <= 11) and (f >=19) and (y < 0.6) and (h >= 4)):
        classes.append({rec[cust]: 'Short Tenure - Promising'})
        is_matched = True
    elif ((r <= 62) and (f >=19) and (y >= 1.5)):
        classes.append({rec[cust]: 'Loyal Customers'})
        is_matched = True
    elif ((r <= 62) and (f >=19) and (y >= 0.6) and (y < 1.5)):
        classes.append({rec[cust]: 'Potential Loyalist'})
        is_matched = True
    elif ((r <= 62) and (f <=18) and (y >= 0.6)):
        classes.append({rec[cust]: 'New Customers'})
        is_matched = True
    
    # 处理未匹配的情况:基于H_cnt找最近的分组
    if not is_matched:
        min_diff = float('inf')
        target_segment = None
        for rule in h_based_rules:
            current_diff = abs(h - rule['threshold'])
            # 更新最小差值对应的分组
            if current_diff < min_diff:
                min_diff = current_diff
                target_segment = rule['segment']
            # 若差值相同,优先选择阈值更高的分组
            elif current_diff == min_diff:
                if rule['threshold'] > [x['threshold'] for x in h_based_rules if abs(h - x['threshold']) == min_diff][0]:
                    target_segment = rule['segment']
        classes.append({rec[cust]: target_segment})

# 统一生成映射,避免循环内重复执行浪费性能
acc_list = [list(item.keys())[0] for item in classes]
segment_list = [list(item.values())[0] for item in classes]
df['new_segment'] = df[cust].map(dict(zip(acc_list, segment_list)))

关键逻辑说明

  1. 预整理H_cnt规则:将所有包含H_cnt阈值的条件提前整理成列表,方便后续计算最近匹配项。
  2. 未匹配处理:计算当前行H_cnt与各规则阈值的绝对差值,找到差值最小的分组;若存在多个差值相同的情况,优先选择阈值更高的分组。
  3. 性能优化:将DataFrame映射逻辑移到循环外,避免重复执行,提升代码效率。
  4. 修正原代码问题:修复了classes_append的拼写错误,以及else分支中无效的字典存储逻辑。

内容的提问来源于stack exchange,提问作者The Great

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 15:04:56