如何基于风速最大值对非规则时间序列气象数据进行重采样?
问题描述
现有一个非规则时间步长的地面风场数据集,包含Time、Speed、Direction等字段(实际数据含更多列及数千行),数据示例如下:
Time, Speed, Direction 2023-1-1 01:00:00, 6, 90 2023-1-1 02:00:00, 6, 70 2023-1-1 03:00:00, 9, 70 2023-1-1 04:00:00, 6, 230 2023-1-1 06:00:00, 2, 320 2023-1-1 08:00:00, 2, 100 2023-1-1 11:00:00, 3, 140 2023-1-1 15:00:00, 12, 10 2023-1-1 16:00:00, 13, 20 2023-1-1 17:00:00, 15, 60 2023-1-1 18:00:00, 10, 80
需要将数据重采样为规则时间步长(00、03、06、09、12、15、18、21点),要求计算每个时间间隔内的最大Speed,并获取对应该最大Speed的Direction。尝试以下代码但无法运行:
df3h = df.resample('3H').agg({ # 3H Does not work if the time series donot start at 00:00 'Speed':'max' 'Direction': lambda x, x.loc[x.Speed.idxmax(),'Direction'] # This Won't Work! })
解决方案
你的代码存在三个核心问题:
- 字典内
Speed与Direction条目间缺少逗号,导致语法错误; lambda函数定义错误,且agg的lambda无法直接访问DataFrame的列;- 重采样起始点未对齐到整点,默认会从数据第一个时间点开始划分区间。
以下是完整的修正方案:
1. 预处理时间列
先确保Time列为datetime类型并设为索引:
import pandas as pd # 加载数据(替换为你的数据加载方式) df = pd.read_csv('wind_data.csv') # 转换时间列格式 df['Time'] = pd.to_datetime(df['Time']) # 设置时间列为索引 df = df.set_index('Time')
2. 对齐重采样时间区间
使用origin='start_day'参数,确保重采样从当天00:00开始,按3小时划分区间:
# 初始化3小时重采样器,对齐到当天整点 resampler = df.resample('3H', origin='start_day')
3. 聚合获取最大Speed及对应Direction
有两种简洁的实现方式:
方式一:自定义聚合函数
def get_max_speed_dir(group): # 找到分组内Speed最大的行(若有多个最大值,取第一个) max_row = group[group['Speed'] == group['Speed'].max()].iloc[0] return pd.Series({'Speed': max_row['Speed'], 'Direction': max_row['Direction']}) # 应用聚合函数 df3h = resampler.apply(get_max_speed_dir)
方式二:通过索引提取
# 获取每个分组中Speed最大的时间索引 max_speed_indices = resampler['Speed'].idxmax() # 根据索引提取对应数据 df3h = df.loc[max_speed_indices] # 将索引替换为重采样区间的起始时间 df3h.index = resampler.groups.keys()
结果验证
对示例数据处理后,输出结果如下:
| Time | Speed | Direction |
|---|---|---|
| 2023-01-01 00:00:00 | 6 | 90 |
| 2023-01-01 03:00:00 | 9 | 70 |
| 2023-01-01 06:00:00 | 2 | 100 |
| 2023-01-01 09:00:00 | 3 | 140 |
| 2023-01-01 12:00:00 | 15 | 60 |
| 2023-01-01 15:00:00 | 15 | 60 |
| 2023-01-01 18:00:00 | 10 | 80 |
| 2023-01-01 21:00:00 | NaN | NaN |
内容的提问来源于stack exchange,提问作者peteron30
相关产品推荐
相关产品推荐

