如何用Pandas简洁获取夜间时段y值首次低于阈值的最早时间
我需要在每日18:00至次日5:00的时段内,找到y值首次低于指定阈值的timestamp。我编写了如下伪代码: ```python timestamps = list() for _, row in df.iterrows(): found = False current = row['timestamp'] val = row['y'] if current is between 6 PM and 5 AM: if not found and value < threshold: found = True timestamps.append(current)
但这段代码不够简洁且易出错,请问是否有更符合Pandas风格的实现方式?
符合Pandas风格的实现方案
1. 预处理时间列
首先确保timestamp列是datetime类型,这是后续时间操作的基础:
df['timestamp'] = pd.to_datetime(df['timestamp'])
2. 标记目标时段
通过提取小时数,标记哪些记录属于18:00至次日5:00的时段:
# 提取时间的小时部分 df['hour'] = df['timestamp'].dt.hour # 标记是否在目标窗口内:18点及以后 或 5点以前 df['in_target_window'] = (df['hour'] >= 18) | (df['hour'] < 5)
3. 按夜间时段分组
因为18:00到次日5:00属于同一个"夜间周期",需要把这些记录归为同一组:
# 对18点及以后的记录,分组日期取当天;5点以前的取前一天,确保整个夜间时段在同一组 df['group_date'] = df['timestamp'].apply( lambda x: x.date() if x.hour >= 18 else x.date() - pd.Timedelta(days=1) )
4. 筛选并获取首次低于阈值的时间
筛选目标时段内y值低于阈值的记录,再按分组取每组第一条的timestamp:
threshold = 600 # 替换为你的实际阈值 # 筛选符合条件的记录,按group_date分组后取每组第一条的timestamp first_timestamps = df[df['in_target_window'] & (df['y'] < threshold)]\ .groupby('group_date')['timestamp']\ .first()\ .tolist()
完整代码示例
import pandas as pd # 假设df是你的原始数据框 df['timestamp'] = pd.to_datetime(df['timestamp']) df['hour'] = df['timestamp'].dt.hour df['in_target_window'] = (df['hour'] >= 18) | (df['hour'] < 5) # 定义夜间时段的分组日期 df['group_date'] = df['timestamp'].apply( lambda x: x.date() if x.hour >= 18 else x.date() - pd.Timedelta(days=1) ) # 设置阈值并获取结果 threshold = 600 first_timestamps = df[df['in_target_window'] & (df['y'] < threshold)]\ .groupby('group_date')['timestamp']\ .first()\ .tolist()
关键优势
- 矢量化操作:避免了低效的
iterrows()循环,处理大数据集时效率提升明显 - 分组逻辑准确:通过
group_date确保整个夜间时段的记录被归为一组,符合需求 - 代码简洁易维护:用Pandas原生方法实现,逻辑清晰,减少出错概率
内容的提问来源于stack exchange,提问作者Luca
相关产品推荐
相关产品推荐

