在R中提取时间戳指定格式年月并排序及数据列扩展需求
patients3 Dataset in bupaR: Step-by-Step Solution Got it, let's tackle your requirements for the patients3 dataset from the bupaR package. I'll walk through each task with clear R code and explanations so you can replicate this easily:
1. Prep: Load Packages & Data
First, make sure you've got bupaR installed and loaded. We'll use base R functions for date-time handling here (no extra packages needed, though you could swap in lubridate if you prefer):
# Install bupaR if you haven't already if (!require(bupaR)) install.packages("bupaR") library(bupaR) # Load the patients3 dataset data(patients3)
2. Sort by Patient ID
First up, we'll sort the dataset to group all records for the same patient together—this is key for calculating meaningful time differences between their events:
# Sort the data by the `patient` column patients_sorted <- patients3[order(patients3$patient), ]
3. Calculate Time Differences (Seconds, Minutes, Hours)
We'll compute the time gap between consecutive timestamps for each patient (using ave() to group calculations by patient). Note that the first entry for each patient will show NA since there's no prior timestamp to compare against:
# Calculate time difference in seconds patients_sorted$time_diff_secs <- ave(as.numeric(patients_sorted$time), patients_sorted$patient, FUN = function(x) c(NA, diff(x))) # Convert seconds to minutes and hours for readability patients_sorted$time_diff_mins <- patients_sorted$time_diff_secs / 60 patients_sorted$time_diff_hours <- patients_sorted$time_diff_mins / 60
4. Add Formatted Year-Month Column (Your First Requirement)
Next, we'll add a new column with the month-year format you specified (e.g., Jan-2017). To make sure we get English month abbreviations (instead of system-default language), we'll set the locale first:
# Set locale to English for consistent month abbreviations Sys.setlocale("LC_TIME", "English") # Insert the year-month column as the 3rd last column in the dataset patients_sorted <- cbind( patients_sorted[, 1:(ncol(patients_sorted)-3)], # All columns except the last 3 time diff cols year_month = format(patients_sorted$time, "%b-%Y"), # New formatted column patients_sorted[, (ncol(patients_sorted)-2):ncol(patients_sorted)] # The 3 time diff cols )
5. Sort by the Formatted Year-Month (Your Second Requirement)
To sort by the new year_month column, we'll convert the string back to a date (using the first day of the month) to ensure proper chronological order (sorting the string directly works, but converting to a date is more reliable):
# Sort the dataset by year-month (ascending order) patients_sorted_by_month <- patients_sorted[ order(as.Date(paste0("01-", patients_sorted$year_month), "%d-%b-%Y")), ] # Optional: Sort by patient AND year-month (keeps patients grouped while sorting by date) patients_sorted_patient_month <- patients_sorted[ order(patients_sorted$patient, as.Date(paste0("01-", patients_sorted$year_month), "%d-%b-%Y")), ]
Quick Notes:
- If you want to remove the
NAvalues from the time difference columns, you can usena.omit(patients_sorted)—just remember this will drop the first record for each patient (since there's no prior timestamp to compare). - If your system doesn't allow setting the locale to English (rare), you can use
lubridate::month(patients_sorted$time, label = TRUE, abbr = TRUE)instead offormat()to get consistent abbreviations.
内容的提问来源于stack exchange,提问作者Robert J

