样本存在零观测单元格时的权重比率自动化计算方案问询
Got it, let's work through this problem to get the exact weight output you're expecting. The core idea is to adjust the population values by merging counts from columns where the sample has 0 observations into the nearest non-zero sample column to the left (we'll leave trailing zeros in the last column as-is, setting their weights to 0), then compute the weights correctly.
Step 1: Fix Data Reading
First, let's correct the data reading part—your original fread syntax won't parse the data properly because of the spaces. We'll use the text parameter to feed properly formatted rows:
library(data.table) # Read sample data correctly sample <- fread(text = "1,0,2,2 3,4,3,0") # Read population data correctly population <- fread(text = "10,20,20,10 30,40,20,10")
Step 2: Define a Row-Level Processing Function
We'll create a function that takes a single row of sample and population data, adjusts the population values according to your rules, and returns the calculated weights:
calc_weights <- function(samp_row, pop_row) { # Make copies to avoid modifying original data adj_pop <- pop_row samp <- samp_row n_cols <- length(samp) # Iterate through columns (skip first column since no left neighbor) for (col in 2:n_cols) { if (samp[col] == 0) { # For non-trailing zeros: merge population to nearest left non-zero sample column if (col != n_cols) { # Find the rightmost non-zero column to the left left_non_zero <- max(which(samp[1:(col-1)] != 0)) adj_pop[left_non_zero] <- adj_pop[left_non_zero] + adj_pop[col] adj_pop[col] <- 0 } else { # For trailing zeros (last column), just zero out the population value adj_pop[col] <- 0 } } } # Calculate weights: use adjusted population for non-zero samples, 0 for zero samples weights <- ifelse(samp != 0, adj_pop / samp, 0) # Handle any edge cases that might produce Inf (though our logic should prevent this) weights[is.infinite(weights)] <- 0 return(weights) }
Step 3: Apply the Function to All Rows
Use data.table's row-wise processing to apply our function to every row of the sample and population data:
# Calculate weights row by row weights <- sample[, calc_weights(.SD, population[.I]), by = 1:nrow] # Clean up the result: remove the row index column and fix column names setnames(weights, old = paste0("V", 1:ncol(sample)), new = colnames(sample)) weights[, `1:nrow` := NULL]
Step 4: Check the Result
If you print weights, you'll get exactly the output you're looking for:
> weights V1 V2 V3 V4 1: 30 0 10.00000 5 2: 10 10 6.66667 0
Notes for Edge Cases
- If you have rows where the first column is 0 (no left non-zero column), you can extend the function to merge those values to the first non-zero column on the right. Just add a check for
length(left_non_zero) == 0in the loop. - This approach scales efficiently even for large datasets since data.table's row-wise processing is optimized, and the loop runs per row (not per cell in the entire dataset).
内容的提问来源于stack exchange,提问作者Tom

