在R中对DataFrame进行子集筛选:保留列首元素并移除空行的实现方法
Great question! Let's work through this with clean, concise R code that matches your requirements. First, let's confirm the task: we need to retain only the first occurrence (topmost non-empty value) in each category1 to category5 column, then remove any rows where all these category columns are completely empty.
First, let's start with your original DataFrame for reference:
df <- data.frame( Hits = c("Hit1", "Hit2", "Hit3", "Hit4", "Hit5"), category1 = c("a1", "", "b1", "a1", "c1"), category2 = c("", "", "", "", "a2"), category3 = c("a3", "", "b3", "", "a3"), category4 = c("", "", "", "", ""), category5 = c("", "", "a5", "b5", ""), stringsAsFactors = FALSE )
Solution 1: Keep only the very first non-empty value per category column
If your goal is to keep only the topmost non-empty entry in each category column (setting all other values in that column to empty), here's a streamlined solution using dplyr:
library(dplyr) df_processed <- df %>% # Target all category columns, retain only the first non-empty value mutate(across(starts_with("category"), ~ ifelse(row_number() == which(. != "")[1], ., ""))) %>% # Drop rows where every category column is empty filter(if_any(starts_with("category"), ~ . != ""))
Let's break this down:
across(starts_with("category"), ...)lets us apply the same logic to all category columns in one gowhich(. != "")[1]finds the row number of the first non-empty value in the column; we keep that value and set everything else to an empty stringfilter(if_any(...))removes rows where all category columns are empty (that's row 2 in your original data)
The resulting df_processed will look like this:
Hits category1 category2 category3 category4 category5 1 Hit1 a1 a3 3 Hit3 a5 5 Hit5 a2
Solution 2: Keep the first occurrence of each unique value per category column
If you actually meant to retain the first instance of every unique non-empty value in each column (and clear out any duplicate values that appear later), use this variation:
df_processed <- df %>% # Clear out duplicate non-empty values in each category column mutate(across(starts_with("category"), ~ ifelse(duplicated(.) & . != "", "", .))) %>% # Remove rows with all empty category columns filter(if_any(starts_with("category"), ~ . != ""))
Here, duplicated(.) & . != "" identifies non-empty values that have appeared earlier in the column; we set those to empty. The output here would be:
Hits category1 category2 category3 category4 category5 1 Hit1 a1 a3 3 Hit3 b1 b3 a5 4 Hit4 b5 5 Hit5 c1 a2
Both solutions are efficient and readable, using dplyr's vectorized operations to avoid messy loops. Pick the one that aligns with your exact needs!
内容的提问来源于stack exchange,提问作者accibio

