如何在pandas中判定多列排序规则并为其分配固定数值别名
实现方案
核心逻辑
- 逐行对所有列按数值降序排序,得到当前行的列名排列序列
- 预先生成所有列的全排列,按字典序排序后给每个排列分配固定数值别名,保证规则统一不随数据变化
- 将每行的排列序列映射为对应的数值别名即可
完整代码
import pandas as pd from itertools import permutations # -------------- 1. 构造示例数据 -------------- df = pd.DataFrame({ 'col1': [5, 1, 7, 1], 'col2': [2, 2, 1, 4], 'col3': [4, 3, 4, 7], 'colN': [9, 9, 0, 8] }) cols = df.columns.tolist() # -------------- 2. 生成固定排列-ID映射表 -------------- # 生成所有列的全排列,按字典序排序后分配唯一固定ID all_permutations = [tuple(p) for p in permutations(cols)] sorted_permutations = sorted(all_permutations) perm_to_id = {perm: idx + 1 for idx, perm in enumerate(sorted_permutations)} # -------------- 3. 逐行计算排列并映射ID -------------- def get_row_sorted_cols(row): # 降序排序,数值相等时按列名升序稳定排序,保证结果可复现 return tuple(row.sort_values(ascending=False, kind='mergesort').index) # 生成排列列 df['row_permutation'] = df.apply(get_row_sorted_cols, axis=1) # 映射为数值别名 df['order'] = df['row_permutation'].map(perm_to_id) # 可选:生成规则描述字符串 df['rule'] = df['row_permutation'].apply(lambda x: ' >= '.join(x))
效果验证
你提到的3列场景下,生成的映射表完全符合你给出的参考规则:
| 排列规则 | 别名 |
|---|---|
| col1 >= col2 >= col3 | 1 |
| col1 >= col3 >= col2 | 2 |
| col2 >= col1 >= col3 | 3 |
| col2 >= col3 >= col1 | 4 |
| col3 >= col1 >= col2 | 5 |
| col3 >= col2 >= col1 | 6 |
规则一致性说明
别名分配逻辑完全固定:只要列名集合不变,全排列的字典序就固定,对应的别名就不会发生变化;如果后续新增列,只需要重新生成全排列映射表即可,分配规则逻辑保持统一。
内容的提问来源于stack exchange,提问作者Ricardo Carrera
相关产品推荐
相关产品推荐

