You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python按月统计商品ID销量:重复代码优化方案咨询

精简代码方案

核心思路

利用Pandas的时间序列处理和分组聚合能力,替代手动逐月份筛选的重复代码,同时用高效的内置统计方法替换自定义的低效频次计算函数。

精简后代码

import pandas as pd

# 读取数据
train_purchases = pd.read_csv(r"C:\Users\Can\Desktop\dressipi_recsys2022\train_purchases.csv")

# 将date列转换为datetime类型,方便按月份处理
train_purchases['date'] = pd.to_datetime(train_purchases['date'])

# 按月+item_id分组,统计每个商品每月的售出次数
monthly_item_counts = train_purchases.groupby(
    [train_purchases['date'].dt.to_period('M'), 'item_id']
).size().reset_index(name='sales_count')

# 找出每个月售出次数最多的商品
top_monthly_items = monthly_item_counts.loc[
    monthly_item_counts.groupby('date')['sales_count'].idxmax()
]

# 遍历输出结果
for _, row in top_monthly_items.iterrows():
    print(f"{row['item_id']} id sold {row['sales_count']} times in {row['date']}.")

代码优化点说明

  1. 消除重复筛选逻辑:
    • 不再手动编写每个月份的日期判断,通过dt.to_period('M')直接提取月份周期(如2020-01),自动覆盖所有15个月的数据。
  2. 替换低效自定义函数:
    • 原most_frequent函数用List.count(i)逐元素统计,时间复杂度为O(n²),数据量大时性能极差。
    • 改用groupby().size()一次性统计所有商品的月度销量,时间复杂度降为O(n),效率大幅提升。
  3. 自动适配全量数据:
    • 无论数据包含多少个月,代码无需修改即可自动处理,扩展性更强。

补充说明

如果需要保留所有商品的月度销量统计(而非仅top1),直接使用monthly_item_counts即可,它包含了每个月份每个商品的售出次数,方便后续进一步分析。

内容的提问来源于stack exchange,提问作者canucar7

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 06:35:26