如何按组迭代data.table?重叠类别优先级分类问题排查
Hey, let's work through this issue together. I see what you're trying to do—handle overlapping category groups by priority, and your current code is misassigning the out value for row 2 (showing c instead of b). Let's break this down and fix it with efficient vectorized operations in data.table (no slow for loops needed, and we'll keep as much of your original logic as possible).
First, Let's Clarify Your Requirements
From your expected output, I can tell you need:
- For rows where
overlap = 0: SetoutandresttoNA(no overlap to process). - For rows in an overlapping group (
overlap = 1):out= the category with the smallest cumulative Priority (i.e., the smallest Priority from the start of the group up to the current row).rest= all other categories in the cumulative set (from group start to current row), in the order they appeared.
What Was Wrong With Your Original Code?
Your original code was taking the global minimum Priority of the entire group and assigning it to every row in the group. That's why row 2 was getting c (the overall group minimum) instead of b (the minimum up to row 2). We need to track the minimum dynamically as we go through each row in the group.
Corrected Code
Here's the revised version that matches your expected output:
library(data.table) # Your original data (updated Priority to match your result example) overlap <- c(0, 1, 1, 1, 0, 1, NA) Priority <- c(3, 2, 1, 4, 5, 6, 7) category <- c("a","b","c","d","e","f","g") data.dt <- data.table(overlap, Priority, category) data.dt$overlap[nrow(data.dt)] <- 0 # Fix last row overlap as you did # Keep your original group logic (detect trailing edges to define overlap blocks) data.dt[, grp := cumsum(c(TRUE, diff(overlap) < 0))] # Process each group to compute cumulative values data.dt[, `:=`( # Track the smallest Priority from group start to current row cum_min_prio = cummin(Priority), # Build a list of categories from group start to current row cum_cats = lapply(seq_len(.N), function(x) category[1:x]) ), by = grp] # Assign `out`: match the cumulative min Priority to its category data.dt[, out := { sapply(seq_len(.N), function(x) { # Get all Priorities and categories up to current row prios_up_to <- Priority[1:x] cats_up_to <- cum_cats[[x]] # Return the category with the smallest Priority so far cats_up_to[which.min(prios_up_to)] }) }, by = grp] # Assign `rest`: all categories except the current `out`, in order data.dt[, rest := { sapply(seq_len(.N), function(x) { current_out <- out[x] current_cats <- cum_cats[[x]] # Join non-out categories into a string toString(current_cats[current_cats != current_out]) }) }, by = grp] # Set NA for non-overlapping rows data.dt[overlap == 0, `:=`(out = NA, rest = NA)] # Clean up temporary columns data.dt[, c("grp", "cum_min_prio", "cum_cats") := NULL] # View the result data.dt[]
Expected Output
Running this code will give you exactly what you wanted:
overlap Priority category out rest 1: 0 3 a <NA> <NA> 2: 1 2 b b a 3: 1 1 c c a,b 4: 1 4 d c a,b,d 5: 0 5 e <NA> <NA> 6: 1 6 f e f 7: 0 7 g <NA> <NA>
Key Fixes & Improvements
- Cumulative Minimum Tracking: Using
cummin()lets us dynamically track the smallest Priority as we move through each row in the group, instead of using the group's global minimum. - Vectorized Group Operations: All processing happens within
data.table's group-by operations, which are far more efficient than manual for loops. - Preserved Original Group Logic: We kept your original
grpcreation code so overlapping blocks are defined exactly as you intended (including the transition from row 5 to 6). - Dynamic
restCalculation: For each row, we build thereststring from the cumulative category list, excluding only the currentoutvalue.
This approach reuses your core grouping logic while fixing the out assignment issue, and keeps things efficient with data.table's optimized operations.
内容的提问来源于stack exchange,提问作者Ipa

