You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr与data.table标记连续出现超3次的Temperature值为Discard?

Hey there! Let's work through how to count consecutive occurrences of values in the Temperature column and flag rows where the streak goes beyond 3 as "Discard". I'll show you solutions using both dplyr (for tidyverse fans) and data.table (for speed, especially with large datasets).

Using dplyr

First, we'll use dplyr's pipeline syntax to group consecutive identical temperature values, calculate their streak length, and apply the flag. We'll use cumsum() with lag() to create group IDs for each consecutive streak:

library(dplyr)

# Create sample data to test with
df <- tibble(
  Temperature = c(22, 22, 22, 22, 25, 25, 28, 28, 28, 28, 28, 30)
)

# Process the data
df_processed <- df %>%
  # Generate a unique ID for each consecutive streak of the same temperature
  mutate(group_id = cumsum(Temperature != lag(Temperature, default = first(Temperature)))) %>%
  # Group by the streak ID to calculate streak length
  group_by(group_id) %>%
  mutate(
    consecutive_count = n(),  # Count how many times the value repeats consecutively
    status = ifelse(consecutive_count > 3, "Discard", "Keep")  # Flag streaks longer than 3
  ) %>%
  ungroup() %>%
  select(-group_id)  # Remove the group ID if you don't need it

# View the result
print(df_processed)

When you run this, you'll see that all rows in the 4-long 22 streak and 5-long 28 streak get marked as "Discard", while shorter streaks stay as "Keep".

Using data.table

If you're working with large datasets, data.table is often faster and more concise. It has a built-in rleid() function that generates IDs for consecutive value streaks—perfect for this task:

library(data.table)

# Convert the sample data to a data.table (or start with one directly)
dt <- as.data.table(df)

# Process the data in one step
dt_processed <- dt[, 
                   `:=`(
                     consecutive_count = .N,  # .N gives the number of rows in the group
                     status = ifelse(.N > 3, "Discard", "Keep")
                   ), 
                   by = rleid(Temperature)]  # Group by consecutive temperature streaks

# View the result
print(dt_processed)

This does exactly the same thing as the dplyr version, but with less code and better performance on big data. The := operator modifies the table in-place (or creates new columns) without copying the entire dataset, which is a big plus for large tables.

Both solutions will give you the same output—pick the one that fits your workflow best!

内容的提问来源于stack exchange,提问作者longlivebrew

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:48:21