data.table滚动计算:基于后续Spuer值调整Eingriff字段逻辑
Efficient Lookahead Check for data.table Eingriff Column
Got it, let's tackle this problem cleanly and scalably—since you need to handle variable lookahead lengths (like 20+ rows), we can avoid manually chaining lead() calls by leveraging data.table's vectorized shift() function. Here's a solution that works for any lookahead size:
Step 1: Setup Your Data
First, let's confirm your initial data (I'll redefine it here for clarity):
library(data.table) DT <- data.table(Zeit = c(1, 2, 3, 4, 5, 6, 7, 8, 9), Spuer = c(45, 45, 32, 25, 30, 44, 34, 42, 44), Eingriff = c(0, 0, 1, 0, 0, 0, 1, 0, 0))
Step 2: Define Your Lookahead Window
Set how many subsequent rows you want to check—change this to 20 or any number for your real data:
lookahead_n <- 3 # Adjust this to your actual required lookahead length
Step 3: Update Eingriff with Vectorized Logic
We can do this in a single concise operation (no messy temporary columns needed):
DT[, Eingriff := fifelse( Eingriff == 1, # Only target rows where Eingriff was originally 1 # Check if any of the next `lookahead_n` Spuer values are <30 (ignore NAs for end-of-table rows) as.integer(sapply(shift(Spuer, n = 1:lookahead_n, type = "lead"), function(x) any(x < 30, na.rm = TRUE))), 0 # Leave non-1 Eingriff values as 0 )]
Step 4: Verify the Result
Running the code above gives you exactly the output you wanted:
DT # Zeit Spuer Eingriff # 1: 1 45 0 # 2: 2 45 0 # 3: 3 32 1 # 4: 4 25 0 # 5: 5 30 0 # 6: 6 44 0 # 7: 7 34 0 # 8: 8 42 0 # 9: 9 44 0
Why This Works
- Scalability: Just change
lookahead_nto 20 (or any number) and it works instantly—no need to add morelead()calls. - Efficiency:
shift()generates all future values in one go, which is much faster than manuallead()chains for large datasets. - Robustness:
na.rm = TRUEensures rows near the end of the table (where there aren't enough subsequent values) don't break the check—we just ignore missing values.
内容的提问来源于stack exchange,提问作者Bolle
相关产品推荐
相关产品推荐

