如何并行移除big.matrix中仅含单一值的列(保留压缩格式)
big.matrix (Parallelized & Compression-Friendly) Great question—handling massive binary matrices with big.matrix while keeping memory usage in check is tricky, especially when you have millions of columns to process. Let’s break down a solution that’s both parallelized and preserves the compressed storage format of big.matrix:
Step 1: Prepare Dependencies & Parallel Environment
First, make sure you have the necessary packages installed. We’ll use bigmemory for core big.matrix handling, bigstatsr for optimized parallel operations on big matrices, and doParallel for setting up parallel workers:
install.packages(c("bigmemory", "bigstatsr", "doParallel")) library(bigmemory) library(bigstatsr) library(doParallel)
Next, register a parallel cluster to leverage multiple CPU cores. Adjust the number of workers based on your system’s capacity (using half your cores is a safe starting point):
cl <- makeCluster(detectCores() %/% 2) registerDoParallel(cl)
Step 2: Efficiently Identify Columns to Keep
Since your matrix is binary (0s and 1s), we can skip counting unique values entirely—instead, check if a column’s sum is 0 (all 0s) or equal to the number of rows (all 1s). These are the columns we want to remove.
Using big_apply (from bigstatsr) is far more efficient than looping through columns one-by-one: it processes columns in blocks, avoids loading the entire matrix into memory, and works seamlessly with parallelization.
# Get matrix dimensions num_rows <- nrow(m4) num_cols <- ncol(m4) # Parallelized check for columns to keep keep_cols <- big_apply( X = m4, a.FUN = function(mat, col_indices) { # Calculate sum of each block of columns (optimized for big matrices) col_sums <- big_colSums(mat[, col_indices, drop = FALSE]) # Keep columns that are NOT all 0 or all 1 !(col_sums == 0 | col_sums == num_rows) }, ind = 1:num_cols, # Process all columns block.size = 10000 # Adjust based on memory: larger blocks = fewer overheads ) # Get indices of columns to retain selected_cols <- which(keep_cols)
Pro Tip for Even More Memory Savings
If your original big.matrix isn’t using the most compact storage, convert it to raw type (1 byte per element, vs 4 bytes for integer) before processing. This cuts memory/disk usage by 75% for binary data:
# Create a compressed raw-type big.matrix m4_compressed <- big.matrix( nrow = num_rows, ncol = num_cols, type = "raw", backingfile = "compressed_matrix.bin", descriptorfile = "compressed_matrix.desc", binary = TRUE # Enables compression ) # Copy data from original matrix to compressed version big_copy(m4, m4_compressed) # Use m4_compressed for the filtering steps above instead of m4
Step 3: Create a New Compressed big.matrix with Filtered Columns
Instead of modifying the original matrix (which is inefficient for disk-backed big.matrix), create a new compressed matrix that only includes the columns we want to keep:
# Initialize new filtered big.matrix with same compression settings m4_filtered <- big.matrix( nrow = num_rows, ncol = length(selected_cols), type = typeof(m4), # Match original type (or use "raw" for compression) backingfile = "filtered_matrix.bin", descriptorfile = "filtered_matrix.desc", binary = TRUE ) # Parallelized copy of selected columns to new matrix big_copy(m4, m4_filtered, ind.col = selected_cols)
Step 4: Clean Up Parallel Resources
Don’t forget to shut down the parallel cluster to free up system resources:
stopCluster(cl)
Key Notes
- Block Size Adjustment: If you run into memory issues, reduce
block.size(e.g., to 5000). If you have plenty of memory, increase it (e.g., to 50000) to speed up processing. - Disk Storage:
big.matrixstores data on disk, so make sure you have enough space for both the original and filtered matrices (you can delete the original files once you confirm the filtered matrix is correct). - Loading Later: To load your filtered matrix in future sessions, use
attach.big.matrix("filtered_matrix.desc").
内容的提问来源于stack exchange,提问作者Keshav M

