如何在R语言数据框中保留每个ID的首次出现并新增唯一ID标识变量
First, let's start with your sample data frame so everyone can follow along:
df <- structure(list( ID = c("Jerry", "Jerry", "Mary", "Tom"), Score = c(65, 98, 88, 75) ), row.names = c(NA, -4L), class = c("tbl_df", "tbl", "data.frame"))
1. Keep only the first occurrence of each ID
There are a couple of straightforward ways to do this with tidyverse tools (since your data is a tibble, this fits perfectly):
Method 1: Using distinct()
The distinct() function lets you keep unique rows for a specified column, and .keep_all = TRUE ensures we retain all other columns in the row:
library(dplyr) df_first_occurrence <- df %>% distinct(ID, .keep_all = TRUE) # Result: # # A tibble: 3 × 2 # ID Score # <chr> <dbl> # 1 Jerry 65 # 2 Mary 88 # 3 Tom 75
Method 2: Using group_by() + slice_head()
Group by the ID column, then take the first row of each group:
df_first_occurrence <- df %>% group_by(ID) %>% slice_head(n = 1) %>% ungroup() # Don't forget to ungroup if you don't need the grouping anymore
This gives the exact same result as the first method.
2. Add a UniqueID column showing ID only on its first occurrence
We can use row_number() within each group to check if a row is the first occurrence, then populate UniqueID accordingly. Here's how:
df_with_uniqueid <- df %>% group_by(ID) %>% mutate(UniqueID = ifelse(row_number() == 1, ID, NA_character_)) %>% ungroup() # Result: # # A tibble: 4 × 3 # ID Score UniqueID # <chr> <dbl> <chr> # 1 Jerry 65 Jerry # 2 Jerry 98 NA # 3 Mary 88 Mary # 4 Tom 75 Tom
Alternatively, if you prefer not to use grouping, you can use the duplicated() function directly:
df_with_uniqueid <- df %>% mutate(UniqueID = ifelse(!duplicated(ID), ID, NA_character_))
This works because duplicated(ID) returns FALSE for the first occurrence of each ID, so !duplicated(ID) is TRUE exactly when we want to show the ID.
内容的提问来源于stack exchange,提问作者Stataq

