R语言中基于occurrence_hour列创建时间分组新列
解决方案
你可以用cut()函数结合自定义分割点和标签实现需求,以下是具体代码:
方法一:Base R 实现
# 加载示例数据集 BEDATA2 <- structure(list(event_unique_id = c("GO-20141260291", "GO-20141260701", "GO-20141260233", "GO-20141260831", "GO-20141260521"), occurrence_date = structure(c(1388534400, 1388534400, 1388534400, 1388534400, 1388534400), tzone = "UTC", class = c("POSIXct", "POSIXt")), occurrence_yrmn = c("2014-January", "2014-January", "2014-January", "2014-January", "2014-January"), reported_date = structure(c(1388534400, 1388534400, 1388534400, 1388534400, 1388534400), tzone = "UTC", class = c("POSIXct", "POSIXt")), reportedy_rmn = c("2014-January", "2014-January", "2014-January", "2014-January", "2014-January"), location_type = c("Other Commercial / Corporate Places (For Profit, Warehouse, Corp. Bldg", "Commercial Dwelling Unit (Hotel, Motel, B & B, Short Term Rental)", "Other Commercial / Corporate Places (For Profit, Warehouse, Corp. Bldg", "Single Home, House (Attach Garage, Cottage, Mobile)", "Bar / Restaurant" ), premises_type = c("Commercial", "Commercial", "Commercial", "House", "Commercial"), reported_dayofweek = c("Wednesday", "Wednesday", "Wednesday", "Wednesday", "Wednesday"), reported_hour = c(1, 3, 2, 3, 2), occurrence_dayofweek = c("Wednesday", "Wednesday", "Wednesday", "Wednesday", "Wednesday"), occurrence_hour = c(1, 3, 1, 3, 2), MCI = c("Break and Enter", "Break and Enter", "Break and Enter", "Break and Enter", "Break and Enter"), Hood_ID = c("71", "70", "126", "136", "81"), Neighbourhood = c("Cabbagetown-South St.James Town", "South Riverdale", "Dorset Park", "West Hill", "Trinity-Bellwoods" ), Object_Id = c(103, 104, 105, 106, 109)), row.names = c(NA, -5L), class = c("tbl_df", "tbl", "data.frame")) # 创建timegroup列 BEDATA2$timegroup <- cut( x = BEDATA2$occurrence_hour, breaks = c(0, 6, 12, 18, 24), # 定义分割点 labels = c( "00-06(午夜至6点)", "06-12(6点至正午)", "12-18(正午至18点)", "18-00(18点至午夜)" ), right = TRUE # 左闭右开区间,确保6点属于下一个组 )
方法二:Tidyverse(dplyr)实现
如果你习惯用tidyverse语法,可以这样写:
library(dplyr) BEDATA2 <- BEDATA2 %>% mutate(timegroup = cut( occurrence_hour, breaks = c(0, 6, 12, 18, 24), labels = c( "00-06(午夜至6点)", "06-12(6点至正午)", "12-18(正午至18点)", "18-00(18点至午夜)" ), right = TRUE ))
参数说明
breaks = c(0,6,12,18,24):将0-23的小时划分为四个区间:[0,6)、[6,12)、[12,18)、[18,24)labels:为每个区间设置自定义名称,匹配你需要的时间窗口描述right = TRUE:默认参数,确保区间左闭右开,符合时间逻辑(比如6点属于"06-12"组,而非"00-06"组)
结果验证
你的示例数据中occurrence_hour均为1、2、3,运行代码后这些行的timegroup都会被归类为"00-06(午夜至6点)",符合预期。
内容的提问来源于stack exchange,提问作者Amantryingtolearn
相关产品推荐
相关产品推荐

