如何提取pandas DataFrame时间列中两个指定时间范围内的数值
问题原因与解决方法
报错原因
你代码触发KeyError的核心问题是:第一行early_birds_df = france_df['time'].str.replace(':','')选取单列后得到的是Series对象,本身不存在time列,后续调用early_birds_df['time']自然找不到对应索引。
正确实现方案
推荐两种稳定的筛选方式,可按需选择:
方法1:转换为时间类型筛选(最规范,无格式兼容问题)
将time列转换为datetime.time类型后直接做区间比对,代码如下:
import pandas as pd # 把time列转为标准时间类型 france_df['time'] = pd.to_datetime(france_df['time'], format='%H:%M').dt.time # 定义筛选区间 start = pd.to_datetime('1:00', format='%H:%M').time() end = pd.to_datetime('3:10', format='%H:%M').time() # 执行筛选 filtered_df = france_df[(france_df['time'] >= start) & (france_df['time'] <= end)]
按照你提供的示例数据,筛选后会保留索引为3-7的5行数据,符合预期。
如果需要筛选5:00-8:00区间,只需替换start和end的时间参数即可。
方法2:字符串直接比对(无需修改原列类型)
因为HH:MM格式的时间字符串字典序和时间顺序完全一致,补全前导零后可直接比对:
# 统一时间格式为5位,补全前导零(如1:17转为01:17) france_df['formatted_time'] = france_df['time'].str.zfill(5) # 执行筛选 filtered_df = france_df[(france_df['formatted_time'] >= '01:00') & (france_df['formatted_time'] <= '03:10')] # 可选:删除临时生成的格式化列 filtered_df = filtered_df.drop('formatted_time', axis=1)
内容的提问来源于stack exchange,提问作者Azurespot
相关产品推荐
相关产品推荐

