在R语言中基于另一变量的数值顺序创建新变量
Hey there! Let's tackle this problem together—you want to add an integer column that tracks the order of each firm's store expansions, right? Here's how you can do it easily with both tidyverse (dplyr) and base R approaches.
First, let's recap your original data so everyone can follow along:
df <- data.frame( firm = c(rep(1,4), rep(2,4)), date = as.Date(c('2017-01-01', '2017-03-01', '2017-05-01', '2017-06-01', '2017-02-01', '2017-04-01', '2017-05-01', '2017-06-01')), city = c('New York', 'DC', 'New York', 'Atlanta', 'DC', 'DC', 'Chicago', 'Atlanta'), numStores = c(1, 1, 2, 1, 1, 2, 1, 1) )
Option 1: Use dplyr (Tidyverse)
This is the most readable approach for most R users. We'll group by firm, sort by date, then assign a rank based on the expansion date.
Case A: Same date = same order (batch expansions)
Use dense_rank() to give the same order number to all expansions on the same date for a firm—perfect if you want to count same-day expansions as a single batch:
library(dplyr) df_with_order <- df %>% group_by(firm) %>% arrange(date) %>% # Ensure rows are sorted by date first mutate(expansion_order = dense_rank(date)) %>% ungroup() # View the result print(df_with_order)
Output:
# A tibble: 8 × 5 firm date city numStores expansion_order <dbl> <date> <chr> <dbl> <int> 1 1 2017-01-01 New York 1 1 2 1 2017-03-01 DC 1 2 3 1 2017-05-01 New York 2 3 4 1 2017-06-01 Atlanta 1 4 5 2 2017-02-01 DC 1 1 6 2 2017-04-01 DC 2 2 7 2 2017-05-01 Chicago 1 3 8 2 2017-06-01 Atlanta 1 4
Case B: Each row gets a unique order (even same date)
If you want a unique sequential number for every row (regardless of shared dates), swap dense_rank() with row_number():
df_with_row_order <- df %>% group_by(firm) %>% arrange(date) %>% mutate(expansion_order = row_number()) %>% ungroup()
Option 2: Base R Approach
If you prefer not to use external packages, here's a base R solution that matches the batch expansion logic:
# First sort the data by firm and date df_sorted <- df[order(df$firm, df$date), ] # Assign order numbers using ave() and dense_rank() df_sorted$expansion_order <- with(df_sorted, ave(as.integer(date), firm, FUN = function(x) dense_rank(x))) # View the result print(df_sorted)
内容的提问来源于stack exchange,提问作者jhelvy

