You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

样本存在零观测单元格时的权重比率自动化计算方案问询

Solution for Weight Calculation with Zero Sample Observations in data.table

Got it, let's work through this problem to get the exact weight output you're expecting. The core idea is to adjust the population values by merging counts from columns where the sample has 0 observations into the nearest non-zero sample column to the left (we'll leave trailing zeros in the last column as-is, setting their weights to 0), then compute the weights correctly.

Step 1: Fix Data Reading

First, let's correct the data reading part—your original fread syntax won't parse the data properly because of the spaces. We'll use the text parameter to feed properly formatted rows:

library(data.table)

# Read sample data correctly
sample <- fread(text = "1,0,2,2
3,4,3,0")

# Read population data correctly
population <- fread(text = "10,20,20,10
30,40,20,10")

Step 2: Define a Row-Level Processing Function

We'll create a function that takes a single row of sample and population data, adjusts the population values according to your rules, and returns the calculated weights:

calc_weights <- function(samp_row, pop_row) {
  # Make copies to avoid modifying original data
  adj_pop <- pop_row
  samp <- samp_row
  n_cols <- length(samp)
  
  # Iterate through columns (skip first column since no left neighbor)
  for (col in 2:n_cols) {
    if (samp[col] == 0) {
      # For non-trailing zeros: merge population to nearest left non-zero sample column
      if (col != n_cols) {
        # Find the rightmost non-zero column to the left
        left_non_zero <- max(which(samp[1:(col-1)] != 0))
        adj_pop[left_non_zero] <- adj_pop[left_non_zero] + adj_pop[col]
        adj_pop[col] <- 0
      } else {
        # For trailing zeros (last column), just zero out the population value
        adj_pop[col] <- 0
      }
    }
  }
  
  # Calculate weights: use adjusted population for non-zero samples, 0 for zero samples
  weights <- ifelse(samp != 0, adj_pop / samp, 0)
  # Handle any edge cases that might produce Inf (though our logic should prevent this)
  weights[is.infinite(weights)] <- 0
  
  return(weights)
}

Step 3: Apply the Function to All Rows

Use data.table's row-wise processing to apply our function to every row of the sample and population data:

# Calculate weights row by row
weights <- sample[, calc_weights(.SD, population[.I]), by = 1:nrow]

# Clean up the result: remove the row index column and fix column names
setnames(weights, old = paste0("V", 1:ncol(sample)), new = colnames(sample))
weights[, `1:nrow` := NULL]

Step 4: Check the Result

If you print weights, you'll get exactly the output you're looking for:

> weights
   V1 V2       V3 V4
1: 30  0 10.00000  5
2: 10 10  6.66667  0

Notes for Edge Cases

  • If you have rows where the first column is 0 (no left non-zero column), you can extend the function to merge those values to the first non-zero column on the right. Just add a check for length(left_non_zero) == 0 in the loop.
  • This approach scales efficiently even for large datasets since data.table's row-wise processing is optimized, and the loop runs per row (not per cell in the entire dataset).

内容的提问来源于stack exchange,提问作者Tom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 14:32:34