You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas简化月度逐小时数据按小时求均值的实现?

批量计算月度各小时均值的优化方案

核心思路

利用Pandas的时间分组与聚合功能,替代手动指定行号的繁琐操作,自动按小时维度计算月度均值,适配任意天数的月份(28/29/30/31天通用)。

分步实现

  1. 确保时间列格式正确
    如果数据集已有时间字段(如timestamp),先转换为datetime类型:

    import pandas as pd
    df1['timestamp'] = pd.to_datetime(df1['timestamp'])
    

    如果没有时间列,可按数据顺序生成对应时间序列(以2024年1月为例):

    df1['timestamp'] = pd.date_range(start='2024-01-01', periods=len(df1), freq='H')
    
  2. 提取小时维度
    从时间列中提取小时数,作为分组依据:

    df1['hour'] = df1['timestamp'].dt.hour
    
  3. 分组计算小时均值
    按hour字段分组,直接聚合计算均值:

    hourly_mean = df1.groupby('hour').mean()
    

完整示例代码

import pandas as pd

# 加载你的数据集(示例:假设df1是已加载的小时级数据)
# df1 = pd.read_csv('your_data.csv')

# 处理时间列(二选一)
# 已有时间列的情况
df1['timestamp'] = pd.to_datetime(df1['timestamp'])
# 无时间列的情况(以2024年1月为例)
# df1['timestamp'] = pd.date_range(start='2024-01-01', periods=len(df1), freq='H')

# 提取小时
df1['hour'] = df1['timestamp'].dt.hour

# 批量计算各小时月度均值
hourly_mean = df1.groupby('hour').mean()

# 输出结果
print(hourly_mean)

扩展说明

如果需要处理跨月份的数据集,同时按月份+小时分组计算:

monthly_hourly_mean = df1.groupby([df1['timestamp'].dt.month, 'hour']).mean()

内容的提问来源于stack exchange,提问作者Kushagra Gupta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 06:48:15