Python处理datetime列识别早中晚等时段出现ValueError如何解决
问题根源
你已经将timestamp_start和timestamp_end列成功转换为datetime64[ns]类型,完全不需要做「转字符串→用strptime解析回时间对象」这步冗余操作,报错就来自这步多余的格式转换:
- 你把datetime对象转字符串时会携带微秒后缀(比如
.828),用"%Y-%m-%d %H:%M:%S"格式解析时,就会因为剩余的微秒字符抛出unconverted data remains错误 - 就算你补了
%f格式符,也属于完全没必要的性能浪费
解决方案
最简修复(直接适配你现有代码逻辑)
直接调用datetime对象自带的time()方法获取时间部分,不需要做字符串转换:
df_time['pickup_time_period'] = df_time['timestamp_start'].apply(lambda x: get_time_period(x.time())) df_time['dropoff_time_period'] = df_time['timestamp_end'].apply(lambda x: get_time_period(x.time()))
更优实现(避免行遍历,性能提升10倍以上)
直接用pandas内置的dt.hour属性做分箱,不需要自定义函数和apply,分箱规则和你原有的get_time_period逻辑完全一致:
import pandas as pd # 定义时段分箱规则 bins = [0,4,10,16,22,24] labels = ['late-night','morning','mid-day','evening','late-night'] # 批量计算上下车时段 df_time['pickup_time_period'] = pd.cut(df_time['timestamp_start'].dt.hour, bins=bins, labels=labels, include_lowest=True) df_time['dropoff_time_period'] = pd.cut(df_time['timestamp_end'].dt.hour, bins=bins, labels=labels, include_lowest=True)
验证方法
你可以先取单行数据测试逻辑是否符合预期:
# 测试第一行上车时间的时段判断 test_time = df_time.loc[0,'timestamp_start'] print(test_time, get_time_period(test_time.time()))
内容的提问来源于stack exchange,提问作者Joehat
相关产品推荐
相关产品推荐

