You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多数据集共通基因筛选与排序:基于R语言的实现需求

Solution: Identify Common Genes & Restructure Sorted Table

Let's walk through how to extract all genes present across your four datasets, sort them alphabetically, and format the table to show each gene's presence/absence in each column.

Step 1: Load Your Example Data

First, let's confirm we're working with the same dataset:

# Reproducible example data
df1 <- read.table(text = " Data1 Data2 Data3 Data4
gene6 - gene1 gene1
- gene2 - gene2
gene3 - gene3 gene5
gene2 gene4 gene2 gene4
gene4 gene1 gene5 -
gene1 gene3 gene6 gene3", header = TRUE, stringsAsFactors = FALSE)

Step 2: Extract & Sort All Unique Genes

We'll pull every gene that appears in any column (ignoring the "-" placeholders), remove duplicates, and sort them alphabetically:

# Get all unique genes, exclude "-", then sort
all_genes <- sort(unique(unlist(df1[df1 != "-"])))

Step 3: Build the Formatted Result Table

For each sorted gene, we'll check if it exists in each dataset column. If it does, we keep the gene name; if not, we use "-":

# Create the result dataframe
result_df <- data.frame(
  Data1 = sapply(all_genes, function(g) if(g %in% df1$Data1) g else "-"),
  Data2 = sapply(all_genes, function(g) if(g %in% df1$Data2) g else "-"),
  Data3 = sapply(all_genes, function(g) if(g %in% df1$Data3) g else "-"),
  Data4 = sapply(all_genes, function(g) if(g %in% df1$Data4) g else "-"),
  row.names = NULL, check.names = FALSE
)

Step 4: View the Final Output

When you print result_df, you'll get exactly the formatted table you're looking for:

result_df

Output:

Data1 Data2 Data3 Data4
1 gene1 gene1 gene1 gene1
2 gene2 gene2 gene2 gene2
3 gene3 gene3 gene3 gene3
4 gene4 gene4     - gene4
5     -     - gene5 gene5
6 gene6     - gene6     -

内容的提问来源于stack exchange,提问作者Paul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:15:33