如何对logcounts1大矩阵按行应用均值±标准差过滤并筛选异常行?
Got it, let's work through this problem clearly. You've got an 8978×4 matrix logcounts1, plus precomputed row means (Average1) and row standard deviations (SDr1), and you need to:
- Flag rows where at least one value is outside the range
[row_mean - row_sd, row_mean + row_sd] - Create an
alphamarker for these rows - Filter the original matrix to keep only those rows
Here's how to do this efficiently in R:
Step 1: Mark individual outlier values
First, create a boolean matrix where each entry is TRUE if the corresponding value in logcounts1 falls outside its row's mean±SD range. R's matrix broadcasting handles this neatly—since Average1 and SDr1 are vectors with length equal to the number of rows, they'll automatically align with each row of the matrix:
# Boolean matrix: TRUE = value is outside row's mean ± sd outlier_per_value <- logcounts1 < (Average1 - SDr1) | logcounts1 > (Average1 + SDr1)
Step 2: Create the alpha marker
Next, generate your alpha marker. If you want a vector indicating which rows have at least one outlier, use rowSums() to count outliers per row, then check if the count is greater than 0:
# alpha as a boolean vector: TRUE = row has at least one outlier alpha <- rowSums(outlier_per_value) > 0 # If you need alpha to match the dimensions of logcounts1 (repeat the row flag for every column): alpha_matrix <- matrix(alpha, nrow = nrow(logcounts1), ncol = ncol(logcounts1), byrow = FALSE)
Step 3: Filter the original matrix
Finally, use the alpha vector to subset logcounts1 and keep only rows with at least one outlier:
# Filtered matrix containing only rows with at least one outlier filtered_logcounts1 <- logcounts1[alpha, ]
Quick verification check
If you want to confirm the logic works, spot-check a flagged row:
# Pick the first row marked as having an outlier sample_row <- which(alpha)[1] # Print details for validation cat("Row", sample_row, "values:", logcounts1[sample_row, ], "\n") cat("Row mean:", Average1[sample_row], "| Row SD:", SDr1[sample_row], "\n") cat("Outlier positions in row:", which(outlier_per_value[sample_row, ]), "\n")
内容的提问来源于stack exchange,提问作者Fernando Delgado Chaves

