如何在R语言数据框中忽略小数对第二列去重(保留首个值)
Got it, let's tackle this problem step by step. The core issue is that your V2 column has decimal suffixes, so we need to first extract the integer portion of those values, then use that (along with V1) to define duplicates and keep the first occurrence.
First, let's assume your data frame looks like this (I've recreated it for clarity):
df <- data.frame( V1 = c("a", "a", "a", "b", "b"), V2 = c("tr78.3", "tr78.2", "tr79.1", "tr12.2", "tr12.3"), stringsAsFactors = FALSE )
Method 1: Using dplyr (tidyverse approach)
This is a clean, readable way if you're already using the tidyverse:
library(dplyr) # Extract integer part from V2, then deduplicate df_cleaned <- df %>% # Create a temporary column with just the integer part of V2 mutate(v2_integer = sub("tr(\\d+)\\.\\d+", "\\1", V2)) %>% # Keep only the first row for each unique combination of V1 and v2_integer distinct(V1, v2_integer, .keep_all = TRUE) %>% # Remove the temporary column we created select(-v2_integer) # View the result df_cleaned
This will output exactly what you want:
V1 V2 1 a tr78.3 2 a tr79.1 3 b tr12.2
Method 2: Base R (no external packages)
If you prefer not to use dplyr, you can do this with base R functions:
# Extract integer part from V2 df$v2_integer <- sub("tr(\\d+)\\.\\d+", "\\1", df$V2) # Identify rows that are NOT duplicates of V1 + v2_integer keep_rows <- !duplicated(paste(df$V1, df$v2_integer)) # Filter the data frame to keep only those rows, and drop the temporary column df_cleaned <- df[keep_rows, !names(df) %in% "v2_integer"] # View the result df_cleaned
Why your initial distinct(df$V1, df$V2) doesn't work
The reason that approach fails is because the full V2 values (like tr78.3 and tr78.2) are technically different strings. By creating the integer-based grouping column, we're telling R to treat any V2 values under the same V1 with matching integer parts as duplicates.
内容的提问来源于stack exchange,提问作者user3224522

