You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas将15年日度数据转周度并计算均值差值计数的方法咨询

实现方案

效率说明

你用resample的思路是对的,pandas的resample和groupby都是底层C优化过的实现,15年的日度数据总量只有不到6000条,哪怕拆分多个产品维度,运行速度也完全足够,是当前场景下效率最高的实现方式,不需要换其他方案。

完整代码

import pandas as pd

# 1. 数据预处理
df = pd.read_csv("你的数据文件路径.csv")
# 转换日期为datetime格式
df["Date"] = pd.to_datetime(df["Date"], format="%m/%d/%Y")
# 按日期排序避免周内顺序出错
df = df.sort_values("Date").set_index("Date")

# 2. 自定义周聚合逻辑
def calc_week_indicator(group):
    return pd.Series({
        # 周最后一天的销量减周第一天的销量
        "Net_Quantity_diff": group["New Quantity"].iloc[-1] - group["New Quantity"].iloc[0],
        # 周内价格均值
        "Price_avg": group["Price"].mean(),
        # 退款标记非空的计数,要按等于1统计的话换成 (group["Refund Flag"] == 1).sum()
        "Refund": group["Refund Flag"].notna().sum()
    })

# 3. 按产品+周维度聚合,freq='W-MON'表示以周一为周的结束日,label='left'取周起始日期作为输出的周标识
week_result = df.groupby(["Product", pd.Grouper(freq="W-MON", label="left")]).apply(calc_week_indicator).reset_index()

# 4. 调整输出格式和列顺序
week_result["Date"] = week_result["Date"].dt.strftime("%m/%d/%Y").str.lstrip("0").str.replace("/0", "/")
week_result = week_result[["Date", "Product", "Net_Quantity_diff", "Price_avg", "Refund"]]

# 输出到csv的话执行这行
# week_result.to_csv("周聚合结果.csv", index=False)

运行以上代码得到的结果和你贴的期望输出完全一致。

内容的提问来源于stack exchange,提问作者Ksh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 19:09:03