如何在Pandas DataFrame中按10:30-13:00时间范围标记班次并筛选数据?
解决Pandas时间范围筛选及班次标记问题
问题根源
你当前的筛选逻辑错误在于用&连接了10点的分钟条件和小时范围,导致11、12点的行也被强制要求分钟≥30,漏掉了这些小时内分钟不足30的行。正确逻辑应为:10点且分钟≥30,或者11-12点(任意分钟),同时时间小于13点。
方案一:直接利用chat_time列(推荐)
无需拆分小时和分钟,直接对chat_time的时间部分筛选,简洁且不易出错:
- 确保
chat_time为datetime类型(若还未转换):
import pandas as pd all_mes_3['chat_time'] = pd.to_datetime(all_mes_3['chat_time'])
- 定义时间范围并标记班次:
# 设定first班次的时间边界 start_first = pd.to_datetime('10:30:00').time() end_first = pd.to_datetime('13:00:00').time() # 新增shift列,先默认标记为'other',再覆盖符合条件的行 all_mes_3['shift'] = 'other' all_mes_3.loc[(all_mes_3['chat_time'].dt.time >= start_first) & (all_mes_3['chat_time'].dt.time < end_first), 'shift'] = 'first'
方案二:修复小时/分钟列的筛选逻辑
如果要继续使用已有的time_h和time_m列,调整筛选条件即可:
# 正确的筛选逻辑:10点且分钟≥30,或者11-12点(任意分钟) first_condition = ((all_mes_3['time_h'] == 10) & (all_mes_3['time_m'] >= 30)) | \ ((all_mes_3['time_h'] > 10) & (all_mes_3['time_h'] < 13)) # 标记first班次 all_mes_3['shift'] = 'other' all_mes_3.loc[first_condition, 'shift'] = 'first'
后续扩展
如果需要标记更多班次,只需新增对应时间条件,用同样的loc赋值方式即可,示例:
# 标记13:00-17:00为second班次 start_second = pd.to_datetime('13:00:00').time() end_second = pd.to_datetime('17:00:00').time() all_mes_3.loc[(all_mes_3['chat_time'].dt.time >= start_second) & (all_mes_3['chat_time'].dt.time < end_second), 'shift'] = 'second'
内容的提问来源于stack exchange,提问作者wihee
相关产品推荐
相关产品推荐

