如何简化Python中多分支if-elif-else的函数代码?
优化权重分段映射函数的几种简洁写法
你的原始代码逻辑清晰但冗余,这里提供几种更简洁高效的实现方式:
方法1:区间映射表+循环匹配
把分段规则整理成有序列表,遍历匹配区间即可,后续修改规则只需调整列表,维护成本低:
def my_fun(x): # 按区间从小到大排列的规则表 slabs = [ (200, 0.2), (300, 0.3), (400, 0.4), (500, 0.5), (600, 0.6), (700, 0.7), (800, 0.8), (900, 0.9), (1000, 1.0) ] for threshold, value in slabs: if x <= threshold: return value # 超出所有区间时返回默认值 return 1.5 courier_invoice['x_weight_slab'] = courier_invoice['weight by company'].apply(my_fun)
方法2:用bisect模块实现二分查找
当分段数量较多时,二分查找比循环匹配效率更高,适合处理大规模数据:
import bisect def my_fun(x): thresholds = [200, 300, 400, 500, 600, 700, 800, 900, 1000] values = [0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0] # 找到第一个大于x的阈值索引 idx = bisect.bisect_right(thresholds, x) return values[idx] if idx < len(values) else 1.5 courier_invoice['x_weight_slab'] = courier_invoice['weight by company'].apply(my_fun)
方法3:直接用Pandas的cut函数(最适配DataFrame场景)
既然你在处理Pandas数据,pd.cut可以完全替代apply和自定义函数,性能更优且代码极简:
import pandas as pd # 定义区间边界与对应标签 bins = [-float('inf'), 200, 300, 400, 500, 600, 700, 800, 900, 1000, float('inf')] labels = [0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.5] courier_invoice['x_weight_slab'] = pd.cut( courier_invoice['weight by company'], bins=bins, labels=labels, include_lowest=True # 确保边界值(如x=200)被正确划分到对应区间 )
注意:你原始代码中存在一个边界重复问题——801<=x<=900和900<=x<=1000会同时匹配x=900,上述优化写法已修正该问题。
内容的提问来源于stack exchange,提问作者Unicorn
相关产品推荐
相关产品推荐

