如何用R或awk将基因-肿瘤突变数据转换为矩阵格式?
Absolutely! Both R and awk are great choices for transforming your two-column text file into the mutation matrix you need. Let’s break down solutions for each tool, starting with a quick example of what I assume your input looks like (adjust if your formatting differs):
Example input (tab-delimited):
gene1 TumorA,TumorC
gene2 TumorB,TumorC,TumorD
gene3 TumorA
R Solution (Using Tidyverse)
The tidyverse suite makes this reshaping task straightforward. We’ll use dplyr for data manipulation and tidyr for splitting and pivoting the data:
# Install and load tidyverse if you haven't already if (!require(tidyverse)) { install.packages("tidyverse") library(tidyverse) } # Read your input file (update the delimiter if your file uses spaces instead of tabs) gene_tumor_data <- read_delim("your_input_file.txt", delim = "\t", col_names = c("gene", "tumors")) # Split tumor lists into individual rows, then pivot to wide format with 1/0 markers mutation_matrix <- gene_tumor_data %>% # Split the tumors column by commas (adjust the separator if your data uses spaces) separate_rows(tumors, sep = ",\\s*") %>% mutate(mutated = 1) %>% # Reshape to wide format, filling missing values with 0 pivot_wider(names_from = tumors, values_from = mutated, values_fill = 0) # Save the final matrix to a new text file write_delim(mutation_matrix, "mutation_matrix_output.txt", delim = "\t")
Key Adjustments for R:
- If your input file uses spaces instead of tabs to separate columns, change
delim = "\t"todelim = " ". - If your tumor lists are separated by spaces instead of commas, update the
separgument inseparate_rowsto"\\s+". - If your input has a header row, remove the
col_namesargument fromread_delimso R uses the existing headers.
awk Solution (Command-Line)
If you prefer a command-line approach without R dependencies, awk works perfectly. This script will first collect all unique genes and tumors, then print the matrix with 1 for mutations and 0 otherwise:
Create a file named mutation_matrix.awk with this code:
BEGIN { FS = "\t"; # Set input field separator to tab (change to " " if using spaces) OFS = "\t"; # Set output field separator to tab } # Skip header row if your input has one (remove this line if there's no header) NR == 1 { next; } # First pass: collect all genes, tumors, and mark mutations { # Track unique genes genes[$1] = 1; # Split the tumor column into an array (adjust regex if your separator isn't commas) split($2, tumor_array, /,\\s*/); for (i in tumor_array) { tumor = tumor_array[i]; tumor_list[tumor] = 1; # Mark that this gene has a mutation in this tumor gene_has_tumor[$1, tumor] = 1; } } # Second pass: print the matrix END { # Print the header row (gene + all tumors) printf "gene"; for (tumor in tumor_list) { printf "%s%s", OFS, tumor; } printf "\n"; # Print each gene's row with 1/0 values for (gene in genes) { printf "%s", gene; for (tumor in tumor_list) { # Print 1 if mutated, else 0 printf "%s%d", OFS, (gene_has_tumor[gene, tumor] ? 1 : 0); } printf "\n"; } }
Run the script from your terminal with:
awk -f mutation_matrix.awk your_input_file.txt > mutation_matrix_output.txt
Key Adjustments for awk:
- If your input uses spaces instead of tabs, change
FS = "\t"toFS = " ". - If tumors are separated by spaces instead of commas, update the split regex to
/\\s+/. - Remove the
NR == 1 { next; }line if your input file doesn’t have a header row.
内容的提问来源于stack exchange,提问作者Meghan Rudd

