R语言入门:从起止时间列统计最常出现的星期几并修复代码问题
错误原因排查
- 时间类型未转换:
read.csv()默认会将时间字段识别为字符串类型,直接调用format()处理字符串无法提取星期信息,必然报错。 - 逻辑偏差:你要统计的是出现频次最高的星期几(众数),和中位数的计算逻辑完全无关,中位数是用于取数值序列的中间分位值,不适合统计最高频类别。
- 语法错误:
order(data.frame, decreasing = TRUE)中data.frame是R的内置关键字,你没有传入实际要排序的数据集,直接运行会报错。 - 冗余且错误的函数调用:
reduce()需要先加载purrr包才能使用,且你写的by分组逻辑完全不符合当前需求,这部分代码属于冗余错误代码。
正确实现代码
使用R基础包即可完成需求,代码如下:
# 读入数据 ny = read.csv('new_york_city.csv') # 1. 将起止时间列从字符串转换为时间格式 # 如果你的csv时间格式和示例不一致,可以调整format参数,运行?strptime可查看所有时间格式标识 ny$Start.Time <- as.POSIXct(ny$Start.Time, format = "%Y-%m-%d %H:%M:%S") ny$End.Time <- as.POSIXct(ny$End.Time, format = "%Y-%m-%d %H:%M:%S") # 2. 提取星期几并统计频次 start_week_count <- table(format(ny$Start.Time, "%A")) end_week_count <- table(format(ny$End.Time, "%A")) # 3. 取频次最高的星期几 most_freq_start <- names(which.max(start_week_count)) most_freq_end <- names(which.max(end_week_count)) # 输出结果 cat("出发时间最常出现的星期:", most_freq_start, ",频次:", max(start_week_count), "\n") cat("结束时间最常出现的星期:", most_freq_end, ",频次:", max(end_week_count), "\n") # 如需查看所有星期的频次排序,可运行以下代码 sort(start_week_count, decreasing = TRUE) sort(end_week_count, decreasing = TRUE)
补充说明
如果需要统计出发+结束合并后的最高频星期,可以用以下代码实现:
all_time <- c(ny$Start.Time, ny$End.Time) all_week_count <- table(format(all_time, "%A")) most_freq_all <- names(which.max(all_week_count))
内容的提问来源于stack exchange,提问作者Sarah Dorra
相关产品推荐
相关产品推荐

