R语言data.table中多列向量值的匹配合并计算方法问询
There isn't a single built-in function that directly implements your custom matching logic, but we can create a reusable function to process each row by splitting the comma-separated values, finding overlapping pairs, merging them, and combining the results. Here's a step-by-step implementation:
Step 1: Define the Processing Function
First, create a function that takes the neg and pos strings from a single row, processes them according to your rules, and returns the combined string of values:
library(data.table) process_matching_values <- function(neg_str, pos_str) { # Split strings into numeric vectors, handling empty strings neg_vals <- as.numeric(strsplit(neg_str, ",\\s*")[[1]]) pos_vals <- as.numeric(strsplit(pos_str, ",\\s*")[[1]]) # Remove NA values (resulting from empty strings) neg_vals <- neg_vals[!is.na(neg_vals)] pos_vals <- pos_vals[!is.na(pos_vals)] # Handle cases where one column has no values if (length(neg_vals) == 0) return(paste(pos_vals, collapse = ", ")) if (length(pos_vals) == 0) return(paste(neg_vals, collapse = ", ")) # Calculate absolute differences between all pairs of values diff_matrix <- outer(neg_vals, pos_vals, function(x, y) abs(x - y)) # Find all pairs where difference is less than 0.1 matching_pairs <- which(diff_matrix < 0.1, arr.ind = TRUE) if (nrow(matching_pairs) > 0) { # Sort pairs by smallest difference first for greedy matching pair_diffs <- diff_matrix[matching_pairs] sorted_pairs <- matching_pairs[order(pair_diffs), ] # Track used values to avoid duplicate matches used_neg <- logical(length(neg_vals)) used_pos <- logical(length(pos_vals)) merged_values <- c() # Iterate through sorted pairs and merge unmatched values for (i in seq(nrow(sorted_pairs))) { neg_idx <- sorted_pairs[i, 1] pos_idx <- sorted_pairs[i, 2] if (!used_neg[neg_idx] && !used_pos[pos_idx]) { merged_values <- c(merged_values, mean(c(neg_vals[neg_idx], pos_vals[pos_idx]))) used_neg[neg_idx] <- TRUE used_pos[pos_idx] <- TRUE } } # Collect remaining unmatched values remaining_neg <- neg_vals[!used_neg] remaining_pos <- pos_vals[!used_pos] combined_values <- c(merged_values, remaining_neg, remaining_pos) } else { # No matches found, combine all original values combined_values <- c(neg_vals, pos_vals) } # Collapse back to comma-separated string paste(combined_values, collapse = ", ") }
Step 2: Apply the Function to Your data.table
Use mapply to apply the function across each row of your dataset, creating a new column with the combined results:
# Your original dataset dataset <- data.table( C = c("a", "b", "c"), neg = c("5, 7", "9", "3, 4, 5"), pos = c("5.05, 8", "", "2.95, 4.2") ) # Add the combined column dataset[, combined := mapply(process_matching_values, neg, pos)]
Step 3: View the Results
After running the code, your dataset will have a new combined column with the merged values:
print(dataset) # C neg pos combined # 1: a 5, 7 5.05, 8 5.025, 7, 8 # 2: b 9 9 # 3: c 3, 4, 5 2.95, 4.2 2.975, 4, 5, 4.2
Key Notes:
- Greedy Matching: The function sorts pairs by their difference and matches the closest pairs first, ensuring each value is merged at most once.
- Empty Values: Handles cases where
posornegis empty by returning the non-empty values directly. - Flexibility: You can adjust the threshold (0.1) by modifying the condition
diff_matrix < 0.1in the function.
内容的提问来源于stack exchange,提问作者xyz

