You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python pandas时间窗口内分类值重复计数新增列(类似滚动value_counts)

解决方案

直接按口味分组后使用 Pandas 原生的时间滚动窗口计数即可,无需生成独热编码,适配多分类场景,性能表现优异。

步骤1:数据预处理

首先将时间列转换为datetime类型,并设置为索引,时间滚动窗口依赖datetime格式的索引运行。

import pandas as pd

# 构造样本数据,实际使用时替换为读入自己的数据集即可
data = [
    ["2014-01-01 07:21:51", "Apple"],
    ["2014-01-01 10:00:47", "Orange"],
    ["2014-01-01 13:24:58", "Banana"],
    ["2014-01-01 15:05:22", "Strawberry"],
    ["2014-01-01 23:26:55", "Lemon"],
    ["2014-01-02 10:07:15", "Orange"],
    ["2014-01-02 10:57:23", "Banana"],
    ["2014-01-03 06:32:11", "Peach"],
    ["2014-01-03 11:29:02", "Orange"],
    ["2014-01-03 19:07:37", "Coconut"],
    ["2014-01-03 19:39:53", "Mango"],
    ["2014-01-04 00:02:36", "Grape"],
    ["2014-01-04 06:51:53", "Cherry"],
    ["2014-01-04 07:53:01", "Strawberry"],
    ["2014-01-04 08:57:48", "Coconut"],
]
df = pd.DataFrame(data, columns=["call_time", "flavor"])

# 转换时间格式+设置索引
df["call_time"] = pd.to_datetime(df["call_time"])
df = df.set_index("call_time")

步骤2:执行滚动计数

按口味分组后,指定72小时窗口滚动计数,closed="both"参数表示包含窗口的首尾边界,匹配需求中的统计规则。

df["count_72h"] = df.groupby("flavor", sort=False)["flavor"]\
    .rolling(window="72h", closed="both")\
    .count()\
    .reset_index(drop=True, level=0)

# 可选:将时间索引还原为普通列
df = df.reset_index()

输出结果

运行后得到的count_72h列与预期结果完全一致,该方案在分类数超过300的场景下依然可以高效运行,无额外内存开销。


内容的提问来源于stack exchange,提问作者Nathalia prada

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 08:00:03