如何获取按总销售额降序排列的前10个唯一prod_title并存入top_tot_sales?
问题需求
从给定数据集中获取按total_sales降序排列的前10个唯一prod_title,并将结果存储到变量top_tot_sales中。
数据集
ID prod_title total_sales 0 619040 All Veggie Yummies 72.99 1 619041 Ball and String 18.95 2 619042 Cat Cave 28.45 3 619043 Chewie Dental 24.95 4 619044 Chomp-a Plush 60.99 5 619045 Feline Fix Mix 65.99 6 619046 Fetch Blaster 9.95 7 619047 Foozy Mouse 45.99 8 619048 Kitty Climber 35.99 9 619049 Purr Mix 32.99 10 619050 Fetch Blaster 19.90 11 619051 Purr Mix 98.97 12 619052 Cat Cave 56.90 13 619053 Purrfect Puree 54.95 14 619054 Foozy Mouse 91.98 15 619055 Reddy Beddy 21.95 16 619056 Cat Cave 85.83 17 619057 Scratchy Post 48.95 18 619058 Snack-em Fish 15.99 19 619059 Snoozer Essentails 99.95 20 619060 Scratchy Post 48.95 21 619061 Purrfect Puree 219.80 22 619062 Chewie Dental 49.90 23 619063 Reddy Beddy 65.85 24 619064 The New Bone 71.96 25 619065 Reddy Beddy 109.75
已尝试的代码
top_tot_sales = df_cleaned.loc[df_cleaned.groupby('prod_title')['total_sales'].idxmax()] df_cleaned.nlargest(10, 'total_sales') df_cleaned['prod_title'].drop_duplicates() df_cleaned['prod_title'].unique() top_tot_sales = df_cleaned.groupby(['prod_title'])['total_sales'].transform(max) == df_cleaned['total_sales'] print(top_tot_sales) df_cleaned['prod_title'].drop_duplicates() df_cleaned['prod_title'].unique() top_tot_sales = df_cleaned.groupby(['prod_title'])df_cleaned.nlargest(n=10, columns=['total_sales']) print(top_tot_sales) top_tot_sales = df_cleaned.groupby('prod_title')['total_sales'].nlargest(n=10) print(top_tot_sales) top_num_sales = df_cleaned.loc[df_cleaned.groupby('prod_title')['trans_quantity'].idxmax()]
正确解决方案
思路解析
要实现需求需分两步:
- 为每个
prod_title保留其最高total_sales的记录,确保产品唯一 - 按
total_sales降序排序后,提取前10个prod_title
实现代码
# 步骤1:获取每个产品销售额最高的记录 max_sales_per_prod = df_cleaned.loc[df_cleaned.groupby('prod_title')['total_sales'].idxmax()] # 步骤2:降序排序后取前10个产品名称 top_tot_sales = max_sales_per_prod.sort_values('total_sales', ascending=False).head(10)['prod_title'].tolist()
代码说明
groupby('prod_title')['total_sales'].idxmax():获取每个产品组中销售额最高的行索引loc[]:根据索引提取对应行,得到每个产品的最高销售额记录sort_values(ascending=False):按销售额降序排列head(10):筛选前10条记录['prod_title'].tolist():提取产品名称并转为列表(若需保留完整行数据,去掉.tolist()即可)
验证结果
基于给定数据集,执行后top_tot_sales的结果为:
['Purrfect Puree', 'Snoozer Essentails', 'Reddy Beddy', 'Purr Mix', 'Foozy Mouse', 'Cat Cave', 'All Veggie Yummies', 'The New Bone', 'Feline Fix Mix', 'Chomp-a Plush']
内容的提问来源于stack exchange,提问作者Ellendejefferes
相关产品推荐
相关产品推荐

