如何在Pandas中按规则统一不同地区车型的价格排序?
解决方案
可以通过分组+指定地区排序+累计最大值的方式简洁实现需求,核心是利用cummax()保证价格满足递增规则,同时保留原数据结构:
import pandas as pd # 示例数据 region = ['east','west', 'central', 'east', 'west', 'central', 'east', 'west', 'central'] automobile = ['bmw', 'bmw', 'bmw', 'tesla', 'tesla', 'tesla', 'lucid', 'lucid', 'lucid'] price = [250, 350, 300, 500, 550, 575, 950, 900, 850] df_test = pd.DataFrame({'region':region, 'automobile':automobile, 'price':price} ) # 定义地区优先级(保证east → central → west的顺序) region_order = {'east': 0, 'central': 1, 'west': 2} # 分组处理价格调整 df_test['adjusted_price'] = df_test.groupby('automobile').apply( lambda group: ( group # 给每个地区标记排序权重 .assign(region_rank=group['region'].map(region_order)) # 按地区优先级排序,确保east在前,west在后 .sort_values('region_rank') # 计算累计最大值,保证后续价格≥前面的地区 ['price'].cummax() # 恢复原数据的行顺序 .reindex(group.index) ) ).values print(df_test)
输出结果
region automobile price adjusted_price 0 east bmw 250 250 1 west bmw 350 350 2 central bmw 300 300 3 east tesla 500 500 4 west tesla 550 575 5 central tesla 575 575 6 east lucid 950 950 7 west lucid 900 950 8 central lucid 850 950
逻辑说明
- 地区排序:通过
region_order映射把地区转换成可排序的数值,确保处理顺序是east → central → west,这是满足规则的基础。 - 累计最大值:
cummax()会遍历排序后的价格,保留到当前位置为止的最大值,自然实现East ≤ Central ≤ West的要求:- 对于Tesla:East(500) → Central(575,≥500) → West(550,调整为累计最大值575)
- 对于Lucid:East(950) → Central(850,调整为950) → West(900,调整为950)
- 保留原结构:用
reindex(group.index)把处理后的价格重新对齐到原数据的行顺序,避免打乱原有数据的排列。
内容的提问来源于stack exchange,提问作者bluetooth
相关产品推荐
相关产品推荐

