多数据集共通基因筛选与排序:基于R语言的实现需求
Solution: Identify Common Genes & Restructure Sorted Table
Let's walk through how to extract all genes present across your four datasets, sort them alphabetically, and format the table to show each gene's presence/absence in each column.
Step 1: Load Your Example Data
First, let's confirm we're working with the same dataset:
# Reproducible example data df1 <- read.table(text = " Data1 Data2 Data3 Data4 gene6 - gene1 gene1 - gene2 - gene2 gene3 - gene3 gene5 gene2 gene4 gene2 gene4 gene4 gene1 gene5 - gene1 gene3 gene6 gene3", header = TRUE, stringsAsFactors = FALSE)
Step 2: Extract & Sort All Unique Genes
We'll pull every gene that appears in any column (ignoring the "-" placeholders), remove duplicates, and sort them alphabetically:
# Get all unique genes, exclude "-", then sort all_genes <- sort(unique(unlist(df1[df1 != "-"])))
Step 3: Build the Formatted Result Table
For each sorted gene, we'll check if it exists in each dataset column. If it does, we keep the gene name; if not, we use "-":
# Create the result dataframe result_df <- data.frame( Data1 = sapply(all_genes, function(g) if(g %in% df1$Data1) g else "-"), Data2 = sapply(all_genes, function(g) if(g %in% df1$Data2) g else "-"), Data3 = sapply(all_genes, function(g) if(g %in% df1$Data3) g else "-"), Data4 = sapply(all_genes, function(g) if(g %in% df1$Data4) g else "-"), row.names = NULL, check.names = FALSE )
Step 4: View the Final Output
When you print result_df, you'll get exactly the formatted table you're looking for:
result_df
Output:
Data1 Data2 Data3 Data4 1 gene1 gene1 gene1 gene1 2 gene2 gene2 gene2 gene2 3 gene3 gene3 gene3 gene3 4 gene4 gene4 - gene4 5 - - gene5 gene5 6 gene6 - gene6 -
内容的提问来源于stack exchange,提问作者Paul
相关产品推荐
相关产品推荐

