You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于变量优先级筛选相似行并将数据集缩减至固定大小的自动化实现

Automated Similar Row Selection for data.table

Got it, let's build a reusable solution to solve this exact problem. You want to iteratively filter rows based on your variable priority, stopping when adding the next variable would reduce the result count below your n_min threshold. Here's an efficient, easy-to-reuse approach using data.table:

Step 1: The Reusable Function

This function takes your data.table, target row, variable priority list, and minimum row count, then returns the most similar rows matching your logic:

library(data.table)

find_similar_rows <- function(dt, target_row, var_priority, n_min) {
  # Pull values from the target row for priority variables
  target_vals <- dt[target_row, ..var_priority]
  
  # Calculate rounded standard deviations for each priority variable (dataset-wide)
  var_sds <- sapply(var_priority, function(var) round(sd(dt[[var]])))
  
  # Start with the highest-priority variable
  current_vars <- var_priority[1]
  current_filter <- paste0(
    current_vars, " %in% (", target_vals[[current_vars]], " - ", var_sds[current_vars], 
    "):(", target_vals[[current_vars]], " + ", var_sds[current_vars], ")"
  )
  current_results <- dt[eval(parse(text = current_filter))]
  
  # Handle edge case: first variable already doesn't meet n_min
  if (nrow(current_results) < n_min) {
    message(paste("Only", nrow(current_results), "rows match the first variable filter (less than n_min =", n_min, "). Returning these."))
    return(current_results)
  }
  
  # Iterate through remaining variables in priority order
  for (var in var_priority[-1]) {
    # Build new filter with the current variable added
    new_filter <- paste0(
      current_filter, " & ", var, " %in% (", target_vals[[var]], " - ", var_sds[var], 
      "):(", target_vals[[var]], " + ", var_sds[var], ")"
    )
    new_results <- dt[eval(parse(text = new_filter))]
    
    # Check if adding this variable drops us below n_min
    if (nrow(new_results) < n_min) {
      message(paste("Adding variable", var, "reduces results to", nrow(new_results), "(below n_min =", n_min, "). Stopping at previous filter."))
      return(current_results)
    }
    
    # Update current filter and results if we still meet n_min
    current_filter <- new_filter
    current_results <- new_results
  }
  
  # If all variables are applied and we still meet n_min, return the full result
  message("All variables applied, result count meets n_min.")
  return(current_results)
}

Step 2: Test with Your Sample Data

Let's verify this works with your example dataset:

set.seed(55)
dat <- data.table(id=1:50, x=sample(100:200,50,replace=TRUE), y=sample(500:600,50,replace=TRUE), z=sample(900:1000,50,replace=TRUE))

# Define your parameters
target_row <- 5  # The row you want to match
var_priority <- c("x", "y", "z")  # Your variable priority order
n_min <- 10  # Minimum number of similar rows needed

# Get the similar rows
similar_rows <- find_similar_rows(dat, target_row, var_priority, n_min)

# Check the result count (should return 16, matching your example)
nrow(similar_rows)

How It Works

  • Target Value Extraction: Pulls the values of your priority variables directly from the target row.
  • Dynamic Interval Calculation: Uses the dataset's standard deviation (rounded) to define the "similar" range for each variable—you can easily swap this for a fixed interval if needed.
  • Iterative Filtering: Starts with your highest-priority variable, then adds each subsequent variable only if the result count stays above n_min.
  • Clear Feedback: Gives simple messages explaining why it stopped at a certain filter step.

Quick Customization Tips

  • Adjust Interval Logic: Replace round(sd(dt[[var]])) with a fixed number (e.g., 15) if you want a consistent ±15 range instead of using standard deviation.
  • Non-Integer Variables: For numeric (non-integer) columns, swap %in% for >= lower & <= upper to avoid missing values due to integer range limits.
  • Sort by Similarity: To get the most similar rows within the filtered set, add a sort step:
    # Add this before returning current_results
    current_results[, similarity_score := abs(x - target_vals$x) + abs(y - target_vals$y) + abs(z - target_vals$z)]
    current_results <- current_results[order(similarity_score)]
    

内容的提问来源于stack exchange,提问作者bumblebee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:36:20