如何在R的多年事件表格中标记事件及提取变量首次出现?
Hey there! Let's work through your two questions using your sample data—super straightforward once you break it down. First, let's build the sample data frame we'll use:
# Create the sample data frame events <- c('event1', 'event1', 'event1', 'event1', 'event1', 'event2', 'event2', 'event2') years <- c('2000', '2001', '2002', '2003', '2004', '1994', '1995', '1996') variable1 <- c('False', 'False', 'False', 'True', 'True', 'False', 'False', 'True') df <- data.frame(events, years, variable1, stringsAsFactors = FALSE)
To flag the first occurrence of each event, we can use grouping operations. The dplyr package makes this intuitive, but I'll also include a base R alternative if you prefer not to load extra packages.
Using dplyr
library(dplyr) # Add a column to mark the first row of each event group df <- df %>% group_by(events) %>% mutate(is_first_event = row_number() == 1) %>% ungroup() # Check the result print(df)
This will output:
# A tibble: 8 × 4 events years variable1 is_first_event <chr> <chr> <chr> <lgl> 1 event1 2000 False TRUE 2 event1 2001 False FALSE 3 event1 2002 False FALSE 4 event1 2003 True FALSE 5 event1 2004 True FALSE 6 event2 1994 False TRUE 7 event2 1995 False FALSE 8 event2 1996 True FALSE
Using Base R
If you want to stick to base R, use the ave() function to check for the first row in each group:
df$is_first_event_base <- ave(seq_along(df$events), df$events, FUN = function(x) x == min(x))
I'll cover two common scenarios here, since "首次出现" could mean two things: the first instance of the variable in the event group, or the first time the variable hits a specific value (like True in your sample).
Scenario 1: Extract the first instance of the variable (first row of each event)
This pulls the very first row for each event, which gives you the initial value of variable1 for that event:
first_variable_instance <- df %>% group_by(events) %>% slice(1) %>% ungroup() print(first_variable_instance)
Output:
# A tibble: 2 × 4 events years variable1 is_first_event <chr> <chr> <chr> <lgl> 1 event1 2000 False TRUE 2 event2 1994 False TRUE
Scenario 2: Extract the first time the variable hits a specific value (e.g., variable1 = 'True')
If you want the first row where variable1 becomes True for each event, use this:
first_true_instance <- df %>% group_by(events) %>% filter(variable1 == 'True') %>% slice(1) %>% ungroup() print(first_true_instance)
Output:
# A tibble: 2 × 4 events years variable1 is_first_event <chr> <chr> <chr> <lgl> 1 event1 2003 True FALSE 2 event2 1996 True FALSE
Base R Alternative for Scenario 2
# Filter rows where variable1 is True, then keep only the first row per event true_rows <- df[df$variable1 == 'True', ] first_true_base <- true_rows[!duplicated(true_rows$events), ]
内容的提问来源于stack exchange,提问作者newhaus94

