使用pandas查找10年时间序列中高频出现最高观测值的15分钟时段
实现方法
核心逻辑
我们的目标是统计10年中,哪个15分钟时段最常成为当日观测值(obs)最高的时段,因此只需要按天筛选最高值记录后做频次统计即可。
代码实现
import pandas as pd # 新增日期、15分钟时段两个辅助字段 stationData10['date'] = stationData10.index.date stationData10['time_slot'] = stationData10.index.time # 计算每日的最大观测值 daily_max_obs = stationData10.groupby('date')['obs'].max().reset_index(name='max_obs') # 关联回原数据集,筛选出所有每日观测值等于最大值的记录 daily_max_records = stationData10.merge(daily_max_obs, on='date') daily_max_records = daily_max_records[daily_max_records['obs'] == daily_max_records['max_obs']] # 统计各15分钟时段出现最高值的频次,按频次从高到低排序 top_time_slots = daily_max_records['time_slot'].value_counts().reset_index() top_time_slots.columns = ['15分钟时段', '出现最高值的天数']
结果使用
- 输出
top_time_slots的第一行就是你要找的最高频出现最高观测值的15分钟时段 - 如果需要按观测站点(location_id)分别统计,只需要在所有groupby操作中加入
location_id作为分组字段即可,示例:# 按站点+日期分组计算最大值 daily_max_obs = stationData10.groupby(['location_id', 'date'])['obs'].max().reset_index(name='max_obs')
可视化优化(可选)
如果需要把频次结果做成热力图,替换你原有代码的聚合逻辑即可:
# 聚合得到各时段各天的最高值出现标记 heatmap_data = daily_max_records.groupby(['time_slot', 'date']).size().unstack(fill_value=0) heatmap_data.hvplot.heatmap( title='每日最高观测值出现时段分布', x='time_slot', y='date', C='value', width=FIGSIZE[0], height=FIGSIZE[1] )
内容的提问来源于stack exchange,提问作者GeoBeez
相关产品推荐
相关产品推荐

