R处理Divvy数据集计算ride_length时触发字符格式报错如何解决?
报错原因
- 核心问题是
started_at和ended_at两列仍为字符类型,difftime()函数执行时需要将字符自动转换为POSIX时间格式,但字符串的时间格式不符合系统默认的可识别标准,因此触发报错。 - 你之前仅提取了
started_at的日期部分生成新列,没有对原始时间列做完整的带时分秒的时间格式转换,也没有对ended_at列做任何格式转换,导致计算时格式不匹配。
解决步骤
- 首先查看时间列的实际格式,确认字符串结构:
head(all_trips_cleaned$started_at)
- 根据实际的时间字符串格式,将两列统一转换为POSIXct标准时间格式,以下为常见的
月/日/年 时:分:秒格式示例,格式符可根据实际输出调整:
# 格式说明:%m=月份,%d=日期,%Y=4位年份,%H=24小时制小时,%M=分钟,%S=秒 # 如果是12小时制带AM/PM,格式改为 "%m/%d/%Y %I:%M:%S %p" all_trips_cleaned <- all_trips_cleaned %>% mutate( started_at = as.POSIXct(started_at, format = "%m/%d/%Y %H:%M:%S"), ended_at = as.POSIXct(ended_at, format = "%m/%d/%Y %H:%M:%S") )
- 转换完成后再计算行程时长:
# units参数显式指定单位为秒,避免默认单位变动导致结果异常 all_trips_cleaned$ride_length <- as.numeric(difftime(all_trips_cleaned$ended_at, all_trips_cleaned$started_at, units = "secs"))
优化建议
你在使用read_csv()导入数据时,可以直接通过col_types参数指定时间列的类型,避免后续重复转换:
# 示例导入单个文件的写法,其他文件同理 Oct_2020_tripdata <- read_csv("Oct 2020.csv", col_types = cols( started_at = col_datetime(format = "%m/%d/%Y %H:%M:%S"), ended_at = col_datetime(format = "%m/%d/%Y %H:%M:%S") ) )
内容的提问来源于stack exchange,提问作者Ugochukwu Orji
相关产品推荐
相关产品推荐

