R语言秒转分钟实现问题求助(骑行时长分析场景)
R语言新手求助:骑行时长秒转分钟及数据类型优化问题
我是R语言新手,正在做骑行时长分析,目前得到的骑行时长结果都是秒单位,不方便做描述性分析和后续可视化,恳请帮忙解决秒转分钟的问题。
现有统计代码及结果
我用以下代码计算骑行时长的统计指标:
trip_stats <- cyclistic_df %>% group_by(member_casual) %>% summarise(average_ride_length = round((mean(ride_length), 2), # average ride length (total ride time / trips) median_ride_length = round(median(ride_length), 2), # median ride length min_ride_length = round(min(ride_length), 2), # minimum ride length max_ride_length = round(max(ride_length), 2)) # maximum ride length head(trip_stats)
运行结果:
# A tibble: 2 × 5 member_casual average_ride_length median_ride_length min_ride_length max_ride_length <chr> <drtn> <drtn> <drtn> <drtn> 1 casual 22.72 secs 785 secs 1 secs 1922127 secs 2 member 12.19 secs 525 secs 1 secs 89872 secs
单独计算平均时长的代码:
# Average ride length (ride_length): ride_lengt_avg <- round(mean(cyclistic_df$ride_length), 2) print(ride_lengt_avg)
运行结果:
Time difference of 977.28 secs
尝试过的方法及问题
我试过as_hms、format("%M:%S")、minutes()等方法,结果还是秒单位。比如我用以下代码提取小时列是成功的:
# Format time as HH:MM:SS: cyclistic_df$time <- format(as.Date(cyclistic_df$date), "%H:%M:%S") # Create new column for time: cyclistic_df$time <- as_hms((cyclistic_df$started_at)) # Create new column for hour: cyclistic_df$hour <- hour(cyclistic_df$time)
运行结果:
0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 casual 32053 20813 12246 6763 4515 8707 23390 40346 55066 56190 72233 93437 109905 114531 121187 134839 153614 170534 member 25582 15645 8736 5268 6097 26049 81660 151664 180585 120035 109540 129717 149403 147804 148800 183330 249039 297724 18 19 20 21 22 23 casual 148452 111916 81388 69348 61476 44808 member 233303 165938 115161 88668 65505 41332
但没法用类似方法把骑行时长转成分钟,直接除以60的话,结果还是会保留"secs"标识,容易造成误导。
已添加的列及疑问
我已经给数据集添加了以下列:
# Default format is yyyy-mm-dd, use start date: cyclistic_df$date <- as.Date(cyclistic_df$started_at) # the default format is yyyy-mm-dd # Create column for year: cyclistic_df$year <- format(as.Date(cyclistic_df$date), "%Y") # Create column for month: cyclistic_df$month <- format(as.Date(cyclistic_df$date), "%m") # Create column for day: cyclistic_df$day <- format(as.Date(cyclistic_df$date), "%d") # Calculate the day of the week: cyclistic_df$day_of_week <- wday(cyclistic_df$started_at) # Create column for day of week: cyclistic_df$day_of_week <- format(as.Date(cyclistic_df$date), "%A") # wday(cyclistic_df$started_at, label = T, abbr = T) # Format time as HH:MM:SS: cyclistic_df$time <- format(as.Date(cyclistic_df$date), "%H:%M:%S") # Create new column for time: cyclistic_df$time <- as_hms((cyclistic_df$started_at)) # Create new column for hour: cyclistic_df$hour <- hour(cyclistic_df$time) # Calculate & Create ride length column by subtracting ended_at time from started_at time and converted it to minutes: cyclistic_df$ride_length <- as_hms(difftime(cyclistic_df$ended_at, cyclistic_df$started_at))
我觉得year、month、day、hour、ride_length这些列应该设为数值类型,避免后续重复转换,是不是最好不要用format( , "%__")这种方法?
内容的提问来源于stack exchange,提问作者prolabrus
相关产品推荐
相关产品推荐

