在R语言中保留A列值出现次数≥10的行
First, let's break down why your current code isn't working:
Your code subset(table(data$A), table(data$A)>=10, drop=FALSE) operates on the frequency table generated by table(data$A), not your original dataset. This gives you a subset of the counts for values in A that meet the threshold, but it doesn't filter the original rows of data—which is why you're seeing deleted rows reappear and columns go missing later on.
Here are three reliable ways to fix this and keep your full dataset intact:
1. Base R: Use ave() to calculate group sizes inline
This method adds a temporary count column to your data, then filters based on that:
# Add a column showing how many times each A value appears data$A_count <- ave(1:nrow(data), data$A, FUN = length) # Filter rows where A appears ≥10 times filtered_data <- data[data$A_count >= 10, ] # Optional: Remove the temporary count column if you don't need it filtered_data$A_count <- NULL
2. Base R: Precompute valid A values with table()
If you don't want to add a temporary column, first get the list of A values that meet the threshold, then filter the original data:
# Get A values that appear ≥10 times # Note: If A is numeric, wrap names() in as.numeric() to avoid type issues valid_A_values <- names(table(data$A))[table(data$A) >= 10] # Filter original dataset to keep only rows with valid A values filtered_data <- data[data$A %in% valid_A_values, ]
3. Tidyverse (dplyr): Clean, readable syntax
If you're using the tidyverse ecosystem, this is the most intuitive approach:
library(dplyr) filtered_data <- data %>% group_by(A) %>% # Group rows by values in A filter(n() >= 10) %>% # Keep only groups with ≥10 rows ungroup() # Remove grouping (optional but good practice)
All three methods will retain all columns from your original dataset and only keep rows where the corresponding A value appears at least 10 times—so your subsequent aggregation operations should work as expected.
内容的提问来源于stack exchange,提问作者Vinzenz

