如何从DataFrame中获取指定Top3客户的访问高峰时段
解决方案
步骤1:提取排名前三的客户ID
因为mostClientDict是按访问次数降序排列的字典(Python 3.7+字典默认保留插入顺序),直接取前3个键即可:
top_3_clients = list(mostClientDict.keys())[:3]
步骤2:从高峰时段数据框中筛选目标客户
如果clientsHours的索引是client_id(groupby后默认结果的索引为分组字段),直接用.loc索引匹配:
top_3_peak_hours = clientsHours.loc[top_3_clients]
如果client_id是clientsHours的普通列,改用.isin()方法筛选:
top_3_peak_hours = clientsHours[clientsHours['client_id'].isin(top_3_clients)]
完整示例(假设clientsHours索引为client_id)
# 提取前三客户ID top_3_clients = list(mostClientDict.keys())[:3] # 获取对应高峰时段 top_3_peak_hours = clientsHours.loc[top_3_clients]
全程无需for循环,完全借助pandas的向量化操作完成筛选。
内容的提问来源于stack exchange,提问作者Guilherme Londero
相关产品推荐
相关产品推荐

