R语言自迭代实现:移除Excel中Name列反向格式重复行
Got it, let's fix this problem step by step. Here's how you can keep only the original non-reversed rows and delete their mirrored counterparts in R:
Step 1: Install & Load Needed Libraries
First, make sure you have the required packages installed. We'll use tidyverse for data wrangling and readxl to read your Excel file:
install.packages(c("tidyverse", "readxl")) library(tidyverse) library(readxl)
Step 2: Import Your Excel Data
Replace the file path below with the actual path to your Excel file:
df <- read_excel("path/to/your/file.xlsx")
Step 3: Filter Out Reversed Rows
The core idea here is to create a standardized key for each entry in the Name column. For example, both "A,Y - B,X" and "B,X - A,Y" will get the same key ("A,Y - B,X") by sorting the two parts of the string. We then keep only rows where the original Name matches this key (these are your "original format" rows):
df_cleaned <- df %>% # Generate a consistent key for each Name entry mutate(standard_key = map_chr(Name, function(x) { # Split the string into two parts using " - " as the separator parts <- str_split(x, " - ")[[1]] # Sort the parts alphabetically and rejoin them paste(sort(parts), collapse = " - ") })) %>% # Keep only rows where the original Name matches the standardized key filter(Name == standard_key) %>% # Remove the helper column we created select(-standard_key)
Step 4: Save the Cleaned Data (Optional)
If you want to export the cleaned data back to an Excel file, use the writexl package:
install.packages("writexl") library(writexl) write_xlsx(df_cleaned, "path/to/save/cleaned_data.xlsx")
How This Works:
map_chrruns our custom function on every entry in theNamecolumn, producing a character vector of standardized keys.- Sorting the two parts of the
Namestring ensures that both original and reversed entries have identical keys. - Filtering where
Name == standard_keykeeps only the entries where the first part comes before the second alphabetically (your original format), and removes all reversed versions along with their associated data columns.
内容的提问来源于stack exchange,提问作者Joe

