基于变量优先级筛选相似行并将数据集缩减至固定大小的自动化实现
Automated Similar Row Selection for data.table
Got it, let's build a reusable solution to solve this exact problem. You want to iteratively filter rows based on your variable priority, stopping when adding the next variable would reduce the result count below your n_min threshold. Here's an efficient, easy-to-reuse approach using data.table:
Step 1: The Reusable Function
This function takes your data.table, target row, variable priority list, and minimum row count, then returns the most similar rows matching your logic:
library(data.table) find_similar_rows <- function(dt, target_row, var_priority, n_min) { # Pull values from the target row for priority variables target_vals <- dt[target_row, ..var_priority] # Calculate rounded standard deviations for each priority variable (dataset-wide) var_sds <- sapply(var_priority, function(var) round(sd(dt[[var]]))) # Start with the highest-priority variable current_vars <- var_priority[1] current_filter <- paste0( current_vars, " %in% (", target_vals[[current_vars]], " - ", var_sds[current_vars], "):(", target_vals[[current_vars]], " + ", var_sds[current_vars], ")" ) current_results <- dt[eval(parse(text = current_filter))] # Handle edge case: first variable already doesn't meet n_min if (nrow(current_results) < n_min) { message(paste("Only", nrow(current_results), "rows match the first variable filter (less than n_min =", n_min, "). Returning these.")) return(current_results) } # Iterate through remaining variables in priority order for (var in var_priority[-1]) { # Build new filter with the current variable added new_filter <- paste0( current_filter, " & ", var, " %in% (", target_vals[[var]], " - ", var_sds[var], "):(", target_vals[[var]], " + ", var_sds[var], ")" ) new_results <- dt[eval(parse(text = new_filter))] # Check if adding this variable drops us below n_min if (nrow(new_results) < n_min) { message(paste("Adding variable", var, "reduces results to", nrow(new_results), "(below n_min =", n_min, "). Stopping at previous filter.")) return(current_results) } # Update current filter and results if we still meet n_min current_filter <- new_filter current_results <- new_results } # If all variables are applied and we still meet n_min, return the full result message("All variables applied, result count meets n_min.") return(current_results) }
Step 2: Test with Your Sample Data
Let's verify this works with your example dataset:
set.seed(55) dat <- data.table(id=1:50, x=sample(100:200,50,replace=TRUE), y=sample(500:600,50,replace=TRUE), z=sample(900:1000,50,replace=TRUE)) # Define your parameters target_row <- 5 # The row you want to match var_priority <- c("x", "y", "z") # Your variable priority order n_min <- 10 # Minimum number of similar rows needed # Get the similar rows similar_rows <- find_similar_rows(dat, target_row, var_priority, n_min) # Check the result count (should return 16, matching your example) nrow(similar_rows)
How It Works
- Target Value Extraction: Pulls the values of your priority variables directly from the target row.
- Dynamic Interval Calculation: Uses the dataset's standard deviation (rounded) to define the "similar" range for each variable—you can easily swap this for a fixed interval if needed.
- Iterative Filtering: Starts with your highest-priority variable, then adds each subsequent variable only if the result count stays above
n_min. - Clear Feedback: Gives simple messages explaining why it stopped at a certain filter step.
Quick Customization Tips
- Adjust Interval Logic: Replace
round(sd(dt[[var]]))with a fixed number (e.g.,15) if you want a consistent ±15 range instead of using standard deviation. - Non-Integer Variables: For numeric (non-integer) columns, swap
%in%for>= lower & <= upperto avoid missing values due to integer range limits. - Sort by Similarity: To get the most similar rows within the filtered set, add a sort step:
# Add this before returning current_results current_results[, similarity_score := abs(x - target_vals$x) + abs(y - target_vals$y) + abs(z - target_vals$z)] current_results <- current_results[order(similarity_score)]
内容的提问来源于stack exchange,提问作者bumblebee
相关产品推荐
相关产品推荐

