基于有效值数量筛选DataFrame:保留至少3个有效值的行
Hey there! Based on the rows you listed to remove, it looks like you want to keep only rows that have at least 3 non-zero values in the numeric columns (116_1 to 116_4). Here are two simple, reliable methods to get this done in R:
Method 1: Base R (No Extra Packages Needed)
This is a straightforward approach using built-in R functions:
# Calculate how many non-zero values exist in each row (skip the Names column) non_zero_counts <- rowSums(aa[, -1] != 0) # Subset the DataFrame to keep only rows with 3+ non-zero values filtered_df <- aa[non_zero_counts >= 3, ]
This works perfectly for your needs:
- Row 1L only has 1 non-zero value, so it gets dropped
- Rows 10L-13L have 2 or fewer non-zero values, so they're removed
- Rows 16L, 25L, 28L, 30L also fall below the 3-non-zero threshold and get filtered out
Method 2: Tidyverse Style with dplyr
If you prefer the pipe-based syntax of the tidyverse, here's a cleaner way to write the same logic:
library(dplyr) filtered_df <- aa %>% # Check non-zero status for all columns except Names, count per row, then filter filter(rowSums(across(-Names, ~ .x != 0)) >= 3)
This does exactly the same work as the base R method, but uses a more readable chain if you're already working with tidyverse tools.
Verify the Result
After running either code, you can check the row names of filtered_df—they should match the rows you wanted to retain. For example, rows like 3L (DDSAI2) and 6L (ADDA1) have all 4 values non-zero, so they stay in the dataset.
内容的提问来源于stack exchange,提问作者Shaxi Liver

