基于data.table的带条件While循环:手机电池衰退率计算
Hey there! Let's figure out how to calculate that battery decay rate with charging cycles properly. Since you're working with a data.table, we'll start with the most efficient approach for that structure, then cover a dplyr alternative, and also fix up your loop attempt.
First, let's confirm the core logic: we group by phone ID, split records into cycles where a new cycle starts when the current battery level is higher than the previous one. For each cycle, we calculate the start/end time, duration, total battery drop, and the decay rate (drop per minute).
Data.table Approach (Most Efficient)
This leverages data.table's fast grouping and vectorized operations, which is perfect for your dataset structure:
library(data.table) # Load your sample data (paste the dput output here if needed) test <- structure(list(time = structure(c(1514229003.91212, 1514229603.61297, 1514230204.14629, 1514230804.81938, 1514231404.36784, 1514232003.73393, 1514232604.17933, 1514233204.00143, 1514233804.68755, 1514234404.15599, 1514235003.99419, 1514235604.68204, 1514236204.18828, 1514236803.66526, 1514237404.0434, 1514238004.40609, 1514238604.02003, 1514239203.42238, 1514239804.19495, 1514240403.15927, 1514241003.87092, 1514241603.93167, 1514242203.77223, 1514242803.66758, 1514243403.33705, 1514244003.25017, 1514244604.05367, 1514245203.7921, 1514245803.2651, 1514246403.63888, 1514247004.02684, 1514247604.04009, 1514248203.99929, 1514248804.07401, 1514249404.11004, 1514250003.74613, 1514250603.88962, 1514251204.19115, 1514251804.06932, 1514252403.94181), class = c("POSIXct", "POSIXt" ), tzone = "EST"), id = c(1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2), level = c(81, 81, 81, 73, 70, 70, 65, 62, 61, 60, 60, 60, 95, 95, 95, 94, 92, 90, 81, 79, 100, 100, 100, 90, 85, 75, 65, 54, 32, 11, 92, 92, 91, 90, 90, 81, 79, 99, 96, 96)), .Names = c("time", "id", "level"), class = c("data.table", "data.frame"), row.names = c(NA, -40L), .internal.selfref = <pointer: 0x102010778>) # Step 1: Assign cycle IDs per phone test[, cycle_id := cumsum(level > shift(level, fill = first(level))), by = id] # Step 2: Calculate cycle metrics outcome <- test[, .( start = min(time), recharge = max(time), difftime = as.numeric(difftime(max(time), min(time), units = "mins")), diffcharge = first(level) - last(level), rate = (first(level) - last(level)) / as.numeric(difftime(max(time), min(time), units = "mins")) ), by = .(id, cycle_id)] # Clean up to match your expected output outcome[, cycle_id := NULL] setcolorder(outcome, c("id", "start", "recharge", "difftime", "diffcharge", "rate")) print(outcome)
How this works:
shift(level, fill = first(level))grabs the previous row's level (we fill the first row with its own level so it doesn't trigger a new cycle).cumsum()turns each "new cycle" flag into a unique ID for every cycle per phone.- We then group by
idandcycle_idto compute all required metrics in one go.
Dplyr Approach
If you prefer the tidyverse syntax, here's an equivalent solution:
library(dplyr) outcome_dplyr <- test %>% group_by(id) %>% # Create cycle IDs (same logic as data.table) mutate(cycle_id = cumsum(level > lag(level, default = first(level)))) %>% group_by(id, cycle_id) %>% summarize( start = min(time), recharge = max(time), difftime = as.numeric(difftime(recharge, start, units = "mins")), diffcharge = first(level) - last(level), rate = diffcharge / difftime ) %>% ungroup() %>% # Clean up columns select(-cycle_id) %>% select(id, start, recharge, difftime, diffcharge, rate) print(outcome_dplyr)
Fixing Your Loop Attempt
Your original loop had a few issues (undefined variables, missing cycle tracking, no boundary handling). Here's a working version using basic loops:
unique_ids <- unique(test$id) outcome_loop <- data.frame() for (current_id in unique_ids) { one_phone <- test[test$id == current_id, ] one_phone <- one_phone[order(one_phone$time), ] if (nrow(one_phone) < 2) next cycle_start_idx <- 1 # Iterate through rows to detect cycle breaks for (i in 2:nrow(one_phone)) { if (one_phone$level[i] > one_phone$level[i-1]) { # Calculate metrics for the completed cycle cycle_data <- one_phone[cycle_start_idx:(i-1), ] cycle_row <- data.frame( id = current_id, start = min(cycle_data$time), recharge = max(cycle_data$time), difftime = as.numeric(difftime(max(cycle_data$time), min(cycle_data$time), units = "mins")), diffcharge = first(cycle_data$level) - last(cycle_data$level), rate = (first(cycle_data$level) - last(cycle_data$level)) / as.numeric(difftime(max(cycle_data$time), min(cycle_data$time), units = "mins")) ) outcome_loop <- rbind(outcome_loop, cycle_row) cycle_start_idx <- i } } # Handle the final cycle that doesn't end with a charge spike cycle_data <- one_phone[cycle_start_idx:nrow(one_phone), ] cycle_row <- data.frame( id = current_id, start = min(cycle_data$time), recharge = max(cycle_data$time), difftime = as.numeric(difftime(max(cycle_data$time), min(cycle_data$time), units = "mins")), diffcharge = first(cycle_data$level) - last(cycle_data$level), rate = (first(cycle_data$level) - last(cycle_data$level)) / as.numeric(difftime(max(cycle_data$time), min(cycle_data$time), units = "mins")) ) outcome_loop <- rbind(outcome_loop, cycle_row) } print(outcome_loop)
Note: For large datasets, the vectorized data.table/dplyr methods will be way faster than loops—stick with those if you can!
内容的提问来源于stack exchange,提问作者iskandarblue

