You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R或awk将基因-肿瘤突变数据转换为矩阵格式?

Absolutely! Both R and awk are great choices for transforming your two-column text file into the mutation matrix you need. Let’s break down solutions for each tool, starting with a quick example of what I assume your input looks like (adjust if your formatting differs):

Example input (tab-delimited):
gene1 TumorA,TumorC
gene2 TumorB,TumorC,TumorD
gene3 TumorA

R Solution (Using Tidyverse)

The tidyverse suite makes this reshaping task straightforward. We’ll use dplyr for data manipulation and tidyr for splitting and pivoting the data:

# Install and load tidyverse if you haven't already
if (!require(tidyverse)) {
  install.packages("tidyverse")
  library(tidyverse)
}

# Read your input file (update the delimiter if your file uses spaces instead of tabs)
gene_tumor_data <- read_delim("your_input_file.txt", 
                              delim = "\t", 
                              col_names = c("gene", "tumors"))

# Split tumor lists into individual rows, then pivot to wide format with 1/0 markers
mutation_matrix <- gene_tumor_data %>%
  # Split the tumors column by commas (adjust the separator if your data uses spaces)
  separate_rows(tumors, sep = ",\\s*") %>%
  mutate(mutated = 1) %>%
  # Reshape to wide format, filling missing values with 0
  pivot_wider(names_from = tumors, 
              values_from = mutated, 
              values_fill = 0)

# Save the final matrix to a new text file
write_delim(mutation_matrix, "mutation_matrix_output.txt", delim = "\t")

Key Adjustments for R:

  • If your input file uses spaces instead of tabs to separate columns, change delim = "\t" to delim = " ".
  • If your tumor lists are separated by spaces instead of commas, update the sep argument in separate_rows to "\\s+".
  • If your input has a header row, remove the col_names argument from read_delim so R uses the existing headers.

awk Solution (Command-Line)

If you prefer a command-line approach without R dependencies, awk works perfectly. This script will first collect all unique genes and tumors, then print the matrix with 1 for mutations and 0 otherwise:

Create a file named mutation_matrix.awk with this code:

BEGIN {
    FS = "\t";  # Set input field separator to tab (change to " " if using spaces)
    OFS = "\t"; # Set output field separator to tab
}

# Skip header row if your input has one (remove this line if there's no header)
NR == 1 { next; }

# First pass: collect all genes, tumors, and mark mutations
{
    # Track unique genes
    genes[$1] = 1;
    # Split the tumor column into an array (adjust regex if your separator isn't commas)
    split($2, tumor_array, /,\\s*/);
    for (i in tumor_array) {
        tumor = tumor_array[i];
        tumor_list[tumor] = 1;
        # Mark that this gene has a mutation in this tumor
        gene_has_tumor[$1, tumor] = 1;
    }
}

# Second pass: print the matrix
END {
    # Print the header row (gene + all tumors)
    printf "gene";
    for (tumor in tumor_list) {
        printf "%s%s", OFS, tumor;
    }
    printf "\n";
    
    # Print each gene's row with 1/0 values
    for (gene in genes) {
        printf "%s", gene;
        for (tumor in tumor_list) {
            # Print 1 if mutated, else 0
            printf "%s%d", OFS, (gene_has_tumor[gene, tumor] ? 1 : 0);
        }
        printf "\n";
    }
}

Run the script from your terminal with:

awk -f mutation_matrix.awk your_input_file.txt > mutation_matrix_output.txt

Key Adjustments for awk:

  • If your input uses spaces instead of tabs, change FS = "\t" to FS = " ".
  • If tumors are separated by spaces instead of commas, update the split regex to /\\s+/.
  • Remove the NR == 1 { next; } line if your input file doesn’t have a header row.

内容的提问来源于stack exchange,提问作者Meghan Rudd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:56:33