You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解析测序比对统计文本文件生成规范DataFrame?

问题:解析测序比对统计文本为结构化DataFrame

原始文本文件格式(重复出现):

Mito_1.fastq.gz
Uniquely mapped reads: 3314106 (74%) 
Multi-mapping reads: 956 (0%) 
Unmapped reads: 1165802 (26%) 
Total: 4480864

Mito_1_old.fastq.gz
Uniquely mapped reads: 1188564 (88%) 
Multi-mapping reads: 406 (0%) 
Unmapped reads: 162676 (12%) 
Total: 1351646

期望输出的DataFrame格式:

> head(desired_outcome)
               Sample Uniquely.mapped.reads Multi.mapping.reads Unmapped.reads   Total
1     Mito_1.fastq.gz               3314106                 956        1165802 4480864
2 Mito_1_old.fastq.gz               1188564                 406         162676 1351646

已尝试的代码:

library(tidyverse)

text_file <- read_tsv("./Results/mapping_stats.txt",
                      col_names = "text")

text_file <- text_file |>
  # Separate numeric values after the colon
  separate(
    col = text,
    into = c("id", "value"),
    sep = "\\:",
    remove = TRUE,
    extra = "warn"
  ) |> 
  # Extract reads and percentage
  mutate(
    reads_number = parse_number(value),
    reads_percentage = str_extract(value, "\\d+(?=%)")
  ) |> 
  # Keep only the relevant columns
  select(id, reads_number)

解决方案

基于你已经完成的步骤,继续以下操作即可得到目标格式:

# 1. 过滤空行(原始文件中的空行会生成id为NA的行)
text_file_clean <- text_file |> filter(!is.na(id))

# 2. 创建样本分组:每次遇到reads_number为NA的行(即样本名行),启动新分组
text_file_clean <- text_file_clean |> 
  mutate(group_id = cumsum(is.na(reads_number)))

# 3. 提取样本名并填充到对应分组的所有行
text_file_clean <- text_file_clean |> 
  group_by(group_id) |> 
  mutate(Sample = first(id[is.na(reads_number)])) |> 
  ungroup()

# 4. 过滤掉样本名行,将统计项转成列(宽格式)
final_df <- text_file_clean |> 
  filter(!is.na(reads_number)) |> 
  pivot_wider(
    id_cols = Sample,
    names_from = id,
    values_from = reads_number
  ) |> 
  # 整理列名,替换空格为点号,匹配期望格式
  rename_with(~str_replace_all(., " ", "."))

# 查看结果
head(final_df)

核心逻辑说明:

  • 用group_id将每个样本的5行数据归为一组
  • 提取每组的样本名,填充到该组所有行
  • 通过pivot_wider将长格式统计项转为宽格式列
  • 最后调整列名格式,与目标输出对齐

内容的提问来源于stack exchange,提问作者neural_axon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 20:47:21