如何为data.table指定列的每个单元格乘以符合特定条件的随机因子
Let's break down how to achieve your goal of multiplying every cell in specified columns of a data.table with a unique random factor (between 0.5-1.5, excluding 1), even when the number of rows is unknown.
Step 1: Create a Robust Factor Generator
First, we need a way to generate valid factors efficiently. Instead of throwing an error if a factor equals 1 (which is statistically extremely unlikely, but possible in edge cases), we'll regenerate any invalid values:
generate_factor_matrix <- function(n_row, n_col) { # Generate initial batch of random factors factor_mat <- matrix(runif(n_row * n_col, min = 0.5, max = 1.5), nrow = n_row, ncol = n_col) # Replace any accidental 1s (edge case handling) while(any(factor_mat == 1)) { invalid_idx <- which(factor_mat == 1) factor_mat[invalid_idx] <- runif(length(invalid_idx), 0.5, 1.5) } return(factor_mat) }
Step 2: Apply the Factors to Your data.table
Using data.table's efficient .SD (Subset of Data) syntax, we can directly multiply the target columns with our factor matrix:
library(data.table) # Your original data.table testDF <- data.table(col1 = c(1,1,1), col2 = c(2,2,2), col3 = c(3,3,3)) # Columns to modify SC <- c("col1", "col3") # Get dimensions of the target subset subset_dims <- dim(testDF[, ..SC]) n_rows <- subset_dims[1] n_cols <- subset_dims[2] # Generate the factor matrix matching the subset dimensions factor_mat <- generate_factor_matrix(n_rows, n_cols) # Perform element-wise multiplication and update the original data.table testDF[, (SC) := .SD * factor_mat, .SDcols = SC]
Why This Works
- Efficiency: Batch-generating random factors is far faster than generating them one-by-one, especially with large datasets.
- data.table Optimization: Using
.SDand.SDcolsleveragesdata.table's optimized column-wise operations, avoiding slow loops. - Robustness: The loop ensures no factor equals 1, covering even the rarest edge cases.
- Flexibility: This works regardless of the number of rows (since we dynamically fetch
n_rowsfrom the subset).
Alternative: Create a Factor data.table
If you prefer working with a data.table instead of a matrix, you can convert the factor matrix and align column names:
factor_dt <- as.data.table(factor_mat) setnames(factor_dt, SC) testDF[, (SC) := .SD * factor_dt, .SDcols = SC]
This achieves the exact same result but might feel more intuitive if you're used to data.table structures.
内容的提问来源于stack exchange,提问作者Nneka

