如何用dplyr与data.table标记连续出现超3次的Temperature值为Discard?
Hey there! Let's work through how to count consecutive occurrences of values in the Temperature column and flag rows where the streak goes beyond 3 as "Discard". I'll show you solutions using both dplyr (for tidyverse fans) and data.table (for speed, especially with large datasets).
First, we'll use dplyr's pipeline syntax to group consecutive identical temperature values, calculate their streak length, and apply the flag. We'll use cumsum() with lag() to create group IDs for each consecutive streak:
library(dplyr) # Create sample data to test with df <- tibble( Temperature = c(22, 22, 22, 22, 25, 25, 28, 28, 28, 28, 28, 30) ) # Process the data df_processed <- df %>% # Generate a unique ID for each consecutive streak of the same temperature mutate(group_id = cumsum(Temperature != lag(Temperature, default = first(Temperature)))) %>% # Group by the streak ID to calculate streak length group_by(group_id) %>% mutate( consecutive_count = n(), # Count how many times the value repeats consecutively status = ifelse(consecutive_count > 3, "Discard", "Keep") # Flag streaks longer than 3 ) %>% ungroup() %>% select(-group_id) # Remove the group ID if you don't need it # View the result print(df_processed)
When you run this, you'll see that all rows in the 4-long 22 streak and 5-long 28 streak get marked as "Discard", while shorter streaks stay as "Keep".
If you're working with large datasets, data.table is often faster and more concise. It has a built-in rleid() function that generates IDs for consecutive value streaks—perfect for this task:
library(data.table) # Convert the sample data to a data.table (or start with one directly) dt <- as.data.table(df) # Process the data in one step dt_processed <- dt[, `:=`( consecutive_count = .N, # .N gives the number of rows in the group status = ifelse(.N > 3, "Discard", "Keep") ), by = rleid(Temperature)] # Group by consecutive temperature streaks # View the result print(dt_processed)
This does exactly the same thing as the dplyr version, but with less code and better performance on big data. The := operator modifies the table in-place (or creates new columns) without copying the entire dataset, which is a big plus for large tables.
Both solutions will give you the same output—pick the one that fits your workflow best!
内容的提问来源于stack exchange,提问作者longlivebrew

