Pandas如何基于Datetime列计算试验总运行时长?KeyError问题如何解决
报错原因
你使用的df[df['Date/Time']]是错误的索引写法:当你向DataFrame的方括号中传入一整列时间序列时,Pandas会将序列中的每个值当作列名去匹配查找,你的数据表中不存在这些时间值命名的列,因此触发KeyError。
报错解决方法
直接对Date/Time列单独调用最大/最小值方法即可,修正后的代码如下:
start = df['Date/Time'].min() end = df['Date/Time'].max()
Pandas的min()、max()方法会自动忽略NaT空值,不需要你额外做过滤处理。
最优总运行时长计算方案
你原本的「最大时间减最小时间」的思路就是最优方案:针对15000行级别的数据量,该操作的时间复杂度为O(n),没有额外性能开销,完整计算代码如下:
# 时间格式转换(你原来的代码无需修改) df['Date/Time'] = pd.to_datetime(df['Date/Time'], format = '%m/%d/%y %H:%M:%S.%f' , errors = 'coerce') # 计算总运行时长(单位:小时,保留2位小数可根据需求调整) total_duration_h = round((df['Date/Time'].max() - df['Date/Time'].min()).total_seconds() / 3600, 2)
如果需要处理全量时间为空的异常场景,可以增加空值判断:
max_time = df['Date/Time'].max() min_time = df['Date/Time'].min() if pd.notna(max_time) and pd.notna(min_time): total_duration_h = round((max_time - min_time).total_seconds() / 3600, 2) else: total_duration_h = 0 # 可替换为你需要的默认值/异常提示逻辑
内容的提问来源于stack exchange,提问作者SeanK22
相关产品推荐
相关产品推荐

