在R语言中查找第n小值对应的列的技术问询
Hey there! Great to hear the initial solution for finding the column with the minimum value worked for you. Let's dive into how to extend this to find columns corresponding to the 2nd, 3rd, ..., nth smallest values. I'll walk through practical scenarios with code examples you can test right away.
Scenario 1: Find columns for the nth smallest value across the entire data frame
First, let's start with a reproducible sample data frame to make things concrete:
# Create a sample data frame with consistent random values set.seed(123) df <- data.frame( ColA = sample(1:10, 5), ColB = sample(1:10, 5), ColC = sample(1:10, 5), ColD = sample(1:10, 5) ) print(df)
To get the column(s) associated with the nth smallest value in the entire dataset (this handles duplicate values too):
n <- 2 # Let's target the 2nd smallest value # Step 1: Flatten the data frame and sort all values sorted_all_vals <- sort(unlist(df)) # Step 2: Grab the nth smallest value target_value <- sorted_all_vals[n] # Step 3: Find all positions matching the target value, then extract column names matching_cols <- colnames(df)[which(df == target_value, arr.ind = TRUE)[, 2]] # Remove duplicates in case the same column has multiple instances of the target value matching_cols <- unique(matching_cols) # Print the result cat(paste("Columns with the", n, "nd smallest value:", paste(matching_cols, collapse = ", ")), "\n")
Scenario 2: Find columns for the nth smallest value per row
This is probably the more common use case—getting the column name for the nth smallest value in each individual row. Here's how to do it:
Basic version (returns first matching column if duplicates exist)
n <- 3 # Target the 3rd smallest value per row # Define a helper function to get the column name for the nth smallest value in a row get_nth_small_col <- function(row, n) { # Order the row values, get the index of the nth smallest, map to column name colnames(df)[order(row)[n]] } # Apply the function to every row row_results <- apply(df, 1, get_nth_small_col, n = n) # Add results to the original data frame for easy viewing df$`3rd_Smallest_Column` <- row_results print(df)
Handle duplicate values (returns all matching columns)
If a row has multiple values equal to the nth smallest, you might want to get all corresponding columns. Adjust the helper function like this:
get_nth_small_cols_with_duplicates <- function(row, n) { sorted_row_vals <- sort(row) target_val <- sorted_row_vals[n] # Return all column names where the row value matches the target colnames(df)[row == target_val] } # Apply to all rows row_results_with_dups <- apply(df, 1, get_nth_small_cols_with_duplicates, n = n) print(row_results_with_dups)
Key Notes
- For the entire data frame approach,
which(df == target_value, arr.ind = TRUE)gives us the row and column indices of all matches—we just extract the column part. - For row-wise operations,
order(row)gives the indices of the sorted row values, soorder(row)[n]points directly to the position of the nth smallest value. - Duplicates: Always consider whether you need just one matching column or all of them, and adjust the code accordingly.
内容的提问来源于stack exchange,提问作者John legend2

