You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何基于另一个DataFrame筛选每个指标对应的top k用户生成结果表

实现方案

核心逻辑:由于仅存在AT1~AT6共6个统计指标,提前预计算每个指标对应的TOP3用户列表,后续直接按客户选择项查表匹配,避免重复计算,完美适配百万级用户规模。


依赖库

仅需要pandas即可完成,无需额外工具:

import pandas as pd

代码实现

1. 读入/构造输入数据

实际场景下替换为你自己的读数据逻辑(比如pd.read_csv读本地文件),下方为样例数据构造代码:

# 构造用户数据表
user_df = pd.DataFrame({
    'user': [1001, 1002, 1003, 1004],
    'AT1': [0.004, 0.2, 0.07, 0.01],
    'AT2': [0.003, 0.1, 0.13, 0.23],
    'AT3': [0.03, 0.3, 0.22, 0.43],
    'AT4': [0.01, 0.1, 0.3, 0.15],
    'AT5': [0.5, 0.1, 0.08, 0.04],
    'AT6': [0.453, 0.2, 0.2, 0.14]
})

# 构造客户选择数据表
client_df = pd.DataFrame({
    'client': [997, 223, 444, 121],
    'choice_1': ['AT2', 'AT6', 'AT1', 'AT1'],
    'choice_2': ['AT3', 'AT5', 'AT4', 'AT5']
})

2. 预计算每个AT指标的TOP3用户字典

# 将用户ID设为索引,对每个AT列降序排序,取前3个用户ID存入字典
top3_per_at = user_df.set_index('user').apply(
    lambda col: col.sort_values(ascending=False).head(3).index.tolist(),
    axis=0
).to_dict()

3. 匹配生成结果表

# 匹配choice_1对应的top1-top3
client_df[['top1', 'top2', 'top3']] = client_df['choice_1'].apply(
    lambda x: pd.Series(top3_per_at[x])
)
# 匹配choice_2对应的top4-top6
client_df[['top4', 'top5', 'top6']] = client_df['choice_2'].apply(
    lambda x: pd.Series(top3_per_at[x])
)

性能说明

  • 全流程仅对6个AT列各做1次全量用户排序,百万级用户可在秒级完成计算
  • 客户匹配阶段仅为O(客户数)的字典查表操作,即使客户量过万也无性能压力

内容的提问来源于stack exchange,提问作者Squid Game

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 07:15:07