如何在R语言中为数据框特定值生成递增序列?
Hey there! Let's tackle this R problem you've got. You want to label each consecutive block of 1s in your var column with sequential numbers (1, 2, 3...) while leaving all 0s as 0. Here are a couple of reliable ways to do this:
Method 1: Base R (No External Packages)
This uses R's built-in rle() (run-length encoding) function, which is perfect for identifying consecutive value blocks:
# Your initial data frame df <- data.frame(var = c(0,0,1,1,0,0,0,1,1,0,0,0,0,1,1)) # Step 1: Get run-length encoding of the var column rle_result <- rle(df$var) # Step 2: Assign sequential numbers to runs of 1s # First, find which runs are 1s, then assign 1,2,3... to them rle_result$values[rle_result$values == 1] <- seq_along(rle_result$values[rle_result$values == 1]) # Step 3: Set all non-1 runs back to 0 rle_result$values[rle_result$values != 1] <- 0 # Step 4: Expand the run-length encoding back to the original length df$newvar <- inverse.rle(rle_result) # View the final result print(df)
How this works:
rle()breaks yourvarcolumn into a list of run lengths and corresponding values (e.g., two 0s, two 1s, three 0s, etc.)- We replace the value of each 1-run with its position in the sequence of 1-runs (first 1-run becomes 1, second becomes 2, etc.)
inverse.rle()converts the modified run-length list back into a full column matching your original data frame's length.
Method 2: Tidyverse (dplyr + Optional data.table)
If you prefer using the tidyverse ecosystem, here's a clean approach. We'll use dplyr for data manipulation, and optionally data.table's rleid() function for simpler consecutive group labeling:
Option 2a: Pure dplyr
library(dplyr) df <- df %>% # Create a unique ID for each consecutive value block mutate(group_id = cumsum(var != lag(var, default = var[1]))) %>% # Assign sequential numbers to 1-blocks, 0 otherwise mutate(newvar = ifelse( var == 1, dense_rank(group_id[var == 1])[match(group_id, group_id[var == 1])], 0 )) %>% # Optional: Remove the temporary group_id column select(-group_id)
Option 2b: dplyr + data.table (Simpler)
rleid() from data.table directly generates IDs for consecutive value blocks, making the code more concise:
library(dplyr) library(data.table) df <- df %>% # Generate consecutive group IDs mutate(group_id = rleid(var)) %>% # Label 1-blocks sequentially, set 0s to 0 mutate(newvar = case_when( var == 1 ~ as.integer(factor(group_id, levels = unique(group_id[var == 1]))), TRUE ~ 0L )) %>% select(-group_id)
Final Result
Both methods will produce this desired data frame:
var newvar 1 0 0 2 0 0 3 1 1 4 1 1 5 0 0 6 0 0 7 0 0 8 1 2 9 1 2 10 0 0 11 0 0 12 0 0 13 0 0 14 1 3 15 1 3
内容的提问来源于stack exchange,提问作者adl

