在R语言中按column3唯一值拆分输出文件并保持格式的方法
在R语言中按指定列的唯一值拆分输出文件
需求说明
现有文件file.txt,内容如下:
column1 column2 column3 column4 column5 column6 column7 column8 column9 column10 column11 column12 column13 column14 column15 column16 1 chr1_10000044_A_T_b38 ENSG00000280113.2 171773 29 30 0.02 0.33 0.144 0.14 chr1 10000044 A T chr1 10060102 2 chr7_10000044_A_T_b38 ENSG00000178585.14 -58627 29 30 0.024 0.26 0.16 0.15 chr7 10000044 A T chr7 18054785 4 chr1_10000044_A_T_b38 ENSG00000280113.2 89708 29 30 0.0 0.03 -0.0 0.038 chr1 10000044 A T chr1 18054638 5 chr1_10000044_A_T_b38 ENSG00000231181.1 -472482 29 30 0.02 0.16 0.11 0.07 chr1 10000044 A T chr1 18052645 6 chr8_304959_A_T_b38 ENSG00000178585.14 -586 60 30 0.026 0.76 0.16 0.15 chr7 10000044 A T chr7 18054785
需要按照column3中的唯一值拆分文件,每个唯一值对应一个输出文件,示例如下:
- 对应
ENSG00000280113.2的文件内容:
column1 column2 column3 column4 column5 column6 column7 column8 column9 column10 column11 column12 column13 column14 column15 column16 1 chr1_10000044_A_T_b38 ENSG00000280113.2 171773 29 30 0.02 0.33 0.144 0.14 chr1 10000044 A T chr1 10060102 4 chr1_10000044_A_T_b38 ENSG00000280113.2 89708 29 30 0.0 0.03 -0.0 0.038 chr1 10000044 A T chr1 18054638
- 对应
ENSG00000178585.14的文件内容:
column1 column2 column3 column4 column5 column6 column7 column8 column9 column10 column11 column12 column13 column14 column15 column16 2 chr7_10000044_A_T_b38 ENSG00000178585.14 -58627 29 30 0.024 0.26 0.16 0.15 chr7 10000044 A T chr7 18054785 6 chr8_304959_A_T_b38 ENSG00000178585.14 -586 60 30 0.026 0.76 0.16 0.15 chr7 10000044 A T chr7 18054785
实现方法
方法1:基础R原生函数实现
无需额外安装包,用原生函数即可完成:
# 读取文件,自动识别任意空白字符作为分隔符 df <- read.table("file.txt", header = TRUE, sep = "", stringsAsFactors = FALSE) # 按column3列分组 split_groups <- split(df, df$column3) # 遍历每个分组,写入对应文件 lapply(names(split_groups), function(group_name) { write.table( split_groups[[group_name]], file = paste0(group_name, ".txt"), # 用分组值作为文件名 sep = " ", # 保持原文件的空格分隔格式 row.names = FALSE, # 不输出行号 quote = FALSE # 不为字符串添加引号 ) })
方法2:tidyverse工具链实现(简洁风格)
如果习惯使用tidyverse系列包,可以用以下代码:
# 首次使用需安装包 # install.packages(c("dplyr", "purrr", "readr")) library(dplyr) library(purrr) library(readr) # 读取文件 df <- read_table("file.txt") # 分组并批量写入文件 df %>% group_split(column3) %>% walk(function(sub_df) { filename <- paste0(unique(sub_df$column3), ".txt") write_delim(sub_df, filename, delim = " ", col_names = TRUE) })
补充说明
- 输出文件名默认用
column3的唯一值命名,比如ENSG00000280113.2.txt,如需自定义格式,可修改paste0(group_name, ".txt")部分,例如改成paste0("result_", group_name, ".txt") - 读取文件时,
read.table(sep="")和read_table()都会自动识别任意空白分隔符,适配原文件格式 - 如果原文件是制表符分隔,可将
sep参数改为"\t"
内容的提问来源于stack exchange,提问作者HKJ3
相关产品推荐
相关产品推荐

