如何批量提取不同目录data.out的7x7矩阵存为CSV供R分析
解决方案
方案1:纯R实现(推荐)
直接在R环境中完成全流程,无需切换工具,输出的CSV可直接用于后续分析,无需依赖固定行号,自动递归处理所有子目录下的data.out文件:
# 依赖包安装(首次运行执行) # install.packages("stringr") library(stringr) # 1. 递归查找所有子目录下的data.out文件 all_data_files <- list.files( path = ".", pattern = "data.out$", recursive = TRUE, full.names = TRUE ) # 2. 定义三类矩阵的匹配关键词 matrix_types <- list( cov = "Estimates of covariance matrix", residual = "Estimates of residual matrix", corr = "Estimates of correlation matrix" ) # 3. 批量处理所有文件 for (file_path in all_data_files) { file_content <- readLines(file_path, warn = FALSE) output_dir <- dirname(file_path) # 逐类提取矩阵 for (mat_type in names(matrix_types)) { # 定位对应矩阵的表头位置 header_idx <- which(str_detect(file_content, matrix_types[[mat_type]])) if (length(header_idx) == 0) next # 跳过表头后的空行、说明行,定位到数据起始行 data_start <- header_idx + 1 while (data_start <= length(file_content) && str_trim(file_content[data_start]) == "") { data_start <- data_start + 1 } data_start <- data_start + 1 # 跳过Matrix/Correlation matrix说明行 # 读取7行矩阵数据 mat_raw_lines <- file_content[data_start:(data_start + 6)] # 解析下三角数据,补全为7x7对称矩阵 full_mat <- matrix(NA, nrow = 7, ncol = 7) for (i in 1:7) { row_values <- as.numeric(str_extract_all(mat_raw_lines[i], "-?\\d+\\.?\\d*")[[1]]) row_values <- row_values[-1] # 去掉每行第一个元素的行号 full_mat[i, 1:i] <- row_values full_mat[1:i, i] <- row_values # 不需要补全上三角可注释掉这行 } # 输出CSV到原data.out所在目录 write.csv( full_mat, file = file.path(output_dir, paste0(mat_type, "_matrix.csv")), row.names = FALSE, col.names = FALSE ) } }
方案2:Shell+awk实现
适合习惯命令行操作的场景,首先将以下代码保存为extract_mat.awk:
#!/usr/bin/awk -f BEGIN { mat_type = "" line_count = 0 } /Estimates of covariance matrix/ {mat_type="cov"; line_count=0; next} /Estimates of residual matrix/ {mat_type="residual"; line_count=0; next} /Estimates of correlation matrix/ {mat_type="corr"; line_count=0; next} mat_type != "" && /^[[:space:]]*[0-9]/ && line_count <7 { for (i=2;i<=NF;i++) { printf "%s", $i if (i<NF) printf "," } printf "\n" line_count++ if (line_count ==7) mat_type="" }
执行以下命令即可批量处理所有data.out文件:
find . -name "data.out" -exec sh -c 'awk -f extract_mat.awk "$1" > "${1%/*}/matrix_output.csv"' _ {} \;
内容的提问来源于stack exchange,提问作者BMT
相关产品推荐
相关产品推荐

