R语言:基于矩阵对应关系计算指定矩阵索引的均值
Got it, let's work through this problem step by step. First, let's fix a small issue in your sample code—you mentioned m1 is a 100×100 matrix, but the original code generates only 1000 elements (which makes a 10×100 matrix). Here's the corrected code to create the 100×100 m1 you described:
# Corrected sample data generation m1 <- matrix(rexp(100*100, rate = .1), ncol = 100) # 100 rows × 100 columns of decimals m2 <- matrix(sample(c(rep(0, 1000), rep(1, 10), rep(2, 1)), 300 * 100, replace = T), 300, 100) # 300 rows × 100 columns of sparse integers
To recap your goal clearly:
- Each row in
m1corresponds to the same-indexed column inm2 - The value at
m2[i,j]represents how many times the j-th row ofm1appears in the i-th observation - We need to calculate a 100-dimensional mean vector for each observation (each row in
m2), where each dimension's mean is the weighted average of the corresponding values fromm1, weighted by their occurrence counts inm2.
Step-by-Step Implementation
Calculate Weighted Sums
Transposem1and multiply it withm2—this gives us a 300×100 matrix where each entry[i,k]is the total weighted sum of the k-th dimension across allm1rows included in the i-th observation. Matrix multiplication is far more efficient than looping here, especially for large matrices.weighted_sums <- m2 %*% t(m1)Get Total Counts per Observation
UserowSums()to compute the total number of occurrences for each observation (row inm2):total_counts <- rowSums(m2)Compute Weighted Means
Divide each row of the weighted sums matrix by the corresponding total count. We also handle cases where an observation has 0 total counts (to avoid division by zero errors):# Use R's broadcasting to divide each row by its total count mean_matrix <- weighted_sums / total_counts # Replace infinite values (from 0/0) with NA (adjust this to 0 or another value if needed) mean_matrix[is.infinite(mean_matrix)] <- NA
What You Get
The resulting mean_matrix is a 300×100 matrix where each row is the mean vector for the corresponding observation in m2. Each entry [i,k] is the average value of the k-th dimension across all m1 rows that appeared in the i-th observation, weighted by how many times each row showed up.
内容的提问来源于stack exchange,提问作者Christopher Costello

