如何在Pandas的else分支中按H_cnt分配行至最近IF条件块
问题描述
现有如下结构的Pandas DataFrame:
customer_id recency frequency H_cnt Years_with_us 1 0 143 3 0.32 2 14 190 8 1.7
核心需求
当某行数据不匹配任何已定义的IF条件时,需基于H_cnt字段将该行分配到最近的IF条件对应的分组。例如:行1的H_cnt=3不匹配任何现有条件,需将其分配至H_cnt>=4对应的「Short Tenure - Promising」分组。
当前代码
cust = 'customer_id' for row in df.iterrows(): rec = row[1] r = rec['recency'] f = rec['frequency'] y = rec['years_with_us'] h = rec['H_cnt'] if ((r <= 11) and (f >=131) and (y >= 0.6) and (h >= 9)): classes.append({rec[cust]:'Champions'}) elif ((r <= 11) and (f >=19) and (y < 0.6) and (h >= 4)): classes.append({rec[cust]:'Short Tenure - Promising'}) elif ((r <= 62) and (f >=19) and (y >= 1.5)): classes.append({rec[cust]:'Loyal Customers'}) elif ((r <= 62) and (f >=19) and (y >= 0.6) and (y < 1.5)): classes.append({rec[cust]:'Potential Loyalist'}) elif ((r <= 62) and (f <=18) and (y >= 0.6)): classes.append({rec[cust]:'New Customers'}) else: print("hi") print(row[1]) classes.append({0:[row[1]['recency'],row[1]['frequency'],row[1]['H_cnt'],row[1]['years_with_IFX']]}) accs = [list(i.keys())[0] for i in classes] segments = [list(i.values())[0] for i in classes] df['new_segment'] = df[cust].map(dict(zip(accs,segments)))
期望输出
customer_id recency frequency H_cnt Years_with_us new_segment 1 0 143 3 0.32 Short Tenure - Promising 2 14 190 8 1.7 Champions
解决方案
修改后的完整代码
cust = 'customer_id' # 预定义带H_cnt阈值的规则及对应分组,按阈值从高到低排序 h_based_rules = [ {'threshold': 9, 'segment': 'Champions'}, {'threshold': 4, 'segment': 'Short Tenure - Promising'} ] classes = [] for row in df.iterrows(): rec = row[1] r = rec['recency'] f = rec['frequency'] y = rec['years_with_us'] h = rec['H_cnt'] is_matched = False # 原有条件判断逻辑 if ((r <= 11) and (f >=131) and (y >= 0.6) and (h >= 9)): classes.append({rec[cust]: 'Champions'}) is_matched = True elif ((r <= 11) and (f >=19) and (y < 0.6) and (h >= 4)): classes.append({rec[cust]: 'Short Tenure - Promising'}) is_matched = True elif ((r <= 62) and (f >=19) and (y >= 1.5)): classes.append({rec[cust]: 'Loyal Customers'}) is_matched = True elif ((r <= 62) and (f >=19) and (y >= 0.6) and (y < 1.5)): classes.append({rec[cust]: 'Potential Loyalist'}) is_matched = True elif ((r <= 62) and (f <=18) and (y >= 0.6)): classes.append({rec[cust]: 'New Customers'}) is_matched = True # 处理未匹配的情况:基于H_cnt找最近的分组 if not is_matched: min_diff = float('inf') target_segment = None for rule in h_based_rules: current_diff = abs(h - rule['threshold']) # 更新最小差值对应的分组 if current_diff < min_diff: min_diff = current_diff target_segment = rule['segment'] # 若差值相同,优先选择阈值更高的分组 elif current_diff == min_diff: if rule['threshold'] > [x['threshold'] for x in h_based_rules if abs(h - x['threshold']) == min_diff][0]: target_segment = rule['segment'] classes.append({rec[cust]: target_segment}) # 统一生成映射,避免循环内重复执行浪费性能 acc_list = [list(item.keys())[0] for item in classes] segment_list = [list(item.values())[0] for item in classes] df['new_segment'] = df[cust].map(dict(zip(acc_list, segment_list)))
关键逻辑说明
- 预整理H_cnt规则:将所有包含
H_cnt阈值的条件提前整理成列表,方便后续计算最近匹配项。 - 未匹配处理:计算当前行
H_cnt与各规则阈值的绝对差值,找到差值最小的分组;若存在多个差值相同的情况,优先选择阈值更高的分组。 - 性能优化:将DataFrame映射逻辑移到循环外,避免重复执行,提升代码效率。
- 修正原代码问题:修复了
classes_append的拼写错误,以及else分支中无效的字典存储逻辑。
内容的提问来源于stack exchange,提问作者The Great
相关产品推荐
相关产品推荐

