如何按SalesPerson分组对Transactiontime升序排名并添加至DataFrame新列
解决方案
步骤说明
- 先将字符串格式的
Transactiontime转换为datetime类型,确保时间排序逻辑正确。 - 按
SalesPerson分组,对组内的Transactiontime进行升序排名,生成整数类型的Rank列。
完整代码
import pandas as pd # 初始化示例数据 dfxx = pd.DataFrame([['2022-02-09 14:00:22', 'A',''], ['2021-10-03 16:58:07', 'B',''], ['2022-01-31 12:24:12', 'X',''], ['2022-01-31 12:24:18', 'X',''], ['2022-01-31 12:24:15', 'X',''], ['2021-10-03 16:58:09', 'B','']], columns=['Transactiontime', 'SalesPerson','Rank']) # 转换时间列为datetime类型 dfxx['Transactiontime'] = pd.to_datetime(dfxx['Transactiontime']) # 分组计算排名 dfxx['Rank'] = dfxx.groupby('SalesPerson')['Transactiontime'].rank(ascending=True, method='first').astype(int) # 输出结果 print(dfxx)
运行结果
| Transactiontime | SalesPerson | Rank | |
|---|---|---|---|
| 0 | 2022-02-09 14:00:22 | A | 1 |
| 1 | 2021-10-03 16:58:07 | B | 1 |
| 2 | 2022-01-31 12:24:12 | X | 1 |
| 3 | 2022-01-31 12:24:18 | X | 3 |
| 4 | 2022-01-31 12:24:15 | X | 2 |
| 5 | 2021-10-03 16:58:09 | B | 2 |
代码细节解释
pd.to_datetime():将字符串时间转为datetime对象,是正确排序和排名的前提。groupby('SalesPerson'):限定排名在每个销售人员的组内独立计算,避免跨组干扰。rank(ascending=True, method='first'):ascending=True:按时间升序排名,更早的交易排名更靠前。method='first':若存在相同时间的交易,按数据中的原始出现顺序分配排名(确保排名唯一)。
astype(int):将排名结果从浮点数转为整数,匹配示例中的格式要求。
内容的提问来源于stack exchange,提问作者j_90
相关产品推荐
相关产品推荐

