R语言:RFID标签监测点停留时长计算的高效优化需求
Hey folks, I'm working with a dataset from an RFID location tracking system where I need to calculate how long each tag stays at a specific monitor point — specifically, the time interval before the MonitorID changes.
Dataset Snippet
I can't generate randomized reproducible data, so here's a partial look at the records:
Time TagID MonitorID Location 2017-10-31 23:03:26 1427435 1352303 A4.18 2017-10-31 23:06:02 1427435 1352303 A4.18 …… 2017-11-22 22:30:55 1427435 1349044 B6.24
Current Function & Performance Bottleneck
I wrote the function below to compute elapsed time per monitor point, and it works as expected functionally. The problem? Processing a single monitor point takes ~4 minutes, and I need to run this for 95 points in batch — total runtime is way too slow for my needs.
Here's the current code:
elapsed_time <- function(x) { # Prepare variables current_monitor <- x$MonitorID[1] start_time <- x$Time[1] end_time <- NULL output <- data.frame("Date" = as.POSIXct(as.character()), "MonitorID" = as.integer(), "Minutes_elapsed" = as.integer()) # For loop to iterate over rows for (i in 1:nrow(x)) { # Skip if monitor hasn't changed (and not last row) if (x$MonitorID[i] == current_monitor & i != nrow(x)) { next } else { # Mark end time when location changes end_time <- x$Time[i] # Calculate time difference in minutes time_spent <- difftime(end_time, start_time, units = "mins") # Create temp row to append temp <- data.frame(start_time, current_monitor, time_spent) # Append to output output <- rbind(output, setNames(temp, names(output))) # Reset start time and current monitor for next segment start_time <- end_time current_monitor <- x$MonitorID[i] } } # Join with monitor-to-location mappings (can ignore this line for optimization focus) output <- left_join(output, Mmappings[,c(1,2)], by="MonitorID") return(output) }
Example Output
Date MonitorID Minutes_elapsed Location 1 2017-10-31 23:03:26 1352303 3.36666667 mins A4.18 …… 6 2017-10-31 23:07:47 0 0.05000000 mins A4.20
(Note: Short intervals between monitor points are due to reading fluctuations, which doesn't impact my core requirements.)
What I'm Looking For
I need practical, efficient optimization strategies to speed up this function, especially for batch processing 95 monitor points. Any ideas to replace the slow loop, improve data handling, or leverage vectorized operations would be a huge help!
内容的提问来源于stack exchange,提问作者anthillcode

