如何检查Pandas DataFrame中时间列是否为15分钟连续间隔
检查15分钟间隔时间序列的连续性
我有一组15分钟间隔的时间序列数据,需要检查数据是否连续。已知该序列缺失2022-02-01 00:45:00+01:00和2022-03-31 23:30:00+01:00两个时间点的数据,示例数据如下:
time consumption 2022-02-01 00:00:00+01:00 107.0 2022-02-01 00:15:00+01:00 57.0 2022-02-01 00:30:00+01:00 177.0 2022-02-01 01:00:00+01:00 30.0 2022-03-31 22:45:00+01:00 21.0 2022-03-31 23:00:00+01:00 25.0 2022-03-31 23:15:00+01:00 43.0 2022-03-31 23:45:00+01:00 30.0
期望输出
- 列出所有缺失的时间点:
"2022-02-01 00:45:00+01:00" is missing "2022-03-31 23:30:00+01:00" is missing
- 返回布尔值判断数据是否连续:
true : 数据完整 false: 至少缺失一个时间点
解决方法(Python Pandas实现)
以下是基于Pandas的具体实现步骤:
1. 数据预处理
将时间列转换为datetime类型并设置为索引:
import pandas as pd # 构造示例DataFrame df = pd.DataFrame({ 'time': ['2022-02-01 00:00:00+01:00', '2022-02-01 00:15:00+01:00', '2022-02-01 00:30:00+01:00', '2022-02-01 01:00:00+01:00', '2022-03-31 22:45:00+01:00', '2022-03-31 23:00:00+01:00', '2022-03-31 23:15:00+01:00', '2022-03-31 23:45:00+01:00'], 'consumption': [107.0, 57.0, 177.0, 30.0, 21.0, 25.0, 43.0, 30.0] }) # 转换时间格式并设置为索引 df['time'] = pd.to_datetime(df['time']) df = df.set_index('time')
2. 生成完整时间序列
根据数据的起止时间,生成15分钟间隔的完整时间序列:
start = df.index.min() end = df.index.max() full_time = pd.date_range(start=start, end=end, freq='15min')
3. 检测缺失时间点
对比完整序列与原始数据,找出缺失的时间点并输出:
missing_times = full_time[~full_time.isin(df.index)] for t in missing_times: print(f'"{t}" is missing')
4. 判断连续性
通过缺失时间点的数量判断数据是否连续:
is_continuous = len(missing_times) == 0 print(f'{is_continuous} : {"数据完整" if is_continuous else "至少缺失一个时间点"}')
执行结果
运行上述代码后,会输出:
"2022-02-01 00:45:00+01:00" is missing "2022-03-31 23:30:00+01:00" is missing False : 至少缺失一个时间点
内容的提问来源于stack exchange,提问作者Naeem
相关产品推荐
相关产品推荐

