人类基因符号转斑马鱼基因符号的R解决方案求助
人类基因名转斑马鱼基因名的R语言解决方案
问题描述
需要将人类(Homo sapiens)基因名称转换为斑马鱼(Danio rerio)基因名称,现有手写的R代码仅能返回每个人类基因对应的最后一个直系同源斑马鱼基因,但实际存在多个同源基因(例如人类基因REG3G对应3个斑马鱼基因:si:ch211-125e6.13、zgc:172053、lectin),希望能输出全部匹配结果。同时需要获取基于BiomaRt的实现方案。
用户原代码如下:
# Read excel file containing list of zebrafish genes and their human orthologs. ortho_genes <- read_excel("/Users/talha/Desktop/Ortho_Gene_List.xlsx") # Separate data from excel file into lists. zebrafish <- ortho_genes$`Zebra Gene Name` human <- ortho_genes$`Human Gene Name` # Read sample list of differential expressed genes sample_list <- c("GREB1L","SIN3B","NCAPG2","FAM50A","PSMD12","BPTF","SLF2","SMC5", "SMC6", "TMEM260","SSBP1","TCF12", "ANLN", "TFAM", "DDX3X","REG3G") # Make a matrix with same number of columns as genes in the supplied list. final_m <- matrix(nrow=length(sample_list),ncol=2) # Iterate through every gene in the supplied list for(x in 1:length(sample_list)){ # Iterate through every human gene for(y in 1:length(human)){ # If the gene from the supplied list matches a human gene if(sample_list[x] == human[y]){ # Fill our matrix in with the supplied gene and the zebrafish ortholog # that matches up with the cell of the human gene final_m[x,1] = sample_list[x] final_m[x,2] = zebrafish[y] } } }
问题原因
原代码使用固定行数的矩阵存储结果,每个人类基因仅对应矩阵的一行,每次匹配到同源基因时会覆盖该行的斑马鱼基因列,最终只保留最后一次匹配的结果,无法容纳多对一的同源关系。
解决方案一:修复现有代码(基于本地同源表)
改用数据框存储结果,每次匹配到同源基因时新增一行,保留所有匹配对:
基础循环版
library(readxl) # 读取本地同源基因表 ortho_genes <- read_excel("/Users/talha/Desktop/Ortho_Gene_List.xlsx") # 差异基因列表 sample_list <- c("GREB1L","SIN3B","NCAPG2","FAM50A","PSMD12","BPTF","SLF2","SMC5", "SMC6", "TMEM260","SSBP1","TCF12", "ANLN", "TFAM", "DDX3X","REG3G") # 初始化空数据框存储结果 final_df <- data.frame(Human_Gene = character(), Zebrafish_Gene = character(), stringsAsFactors = FALSE) # 遍历每个目标人类基因 for(gene in sample_list){ # 筛选出当前基因对应的所有斑马鱼同源基因 matched_rows <- ortho_genes[ortho_genes$`Human Gene Name` == gene, ] # 若存在匹配结果,添加到结果数据框 if(nrow(matched_rows) > 0){ temp_df <- data.frame(Human_Gene = rep(gene, nrow(matched_rows)), Zebrafish_Gene = matched_rows$`Zebra Gene Name`, stringsAsFactors = FALSE) final_df <- rbind(final_df, temp_df) } } # 查看结果 print(final_df)
dplyr简洁版
如果熟悉dplyr包,可直接用过滤选择操作实现,代码更简洁高效:
library(readxl) library(dplyr) ortho_genes <- read_excel("/Users/talha/Desktop/Ortho_Gene_List.xlsx") sample_list <- c("GREB1L","SIN3B","NCAPG2","FAM50A","PSMD12","BPTF","SLF2","SMC5", "SMC6", "TMEM260","SSBP1","TCF12", "ANLN", "TFAM", "DDX3X","REG3G") final_df <- ortho_genes %>% # 筛选出目标人类基因对应的行 filter(`Human Gene Name` %in% sample_list) %>% # 重命名列名并保留需要的列 select(Human_Gene = `Human Gene Name`, Zebrafish_Gene = `Zebra Gene Name`) print(final_df)
解决方案二:基于BiomaRt直接获取同源基因
无需本地同源表,直接通过Ensembl数据库的BiomaRt接口获取跨物种直系同源基因:
library(biomaRt) # 连接Ensembl数据库(国内访问慢可切换镜像:useEnsembl(biomart = "ensembl", mirror = "asia")) ensembl <- useEnsembl(biomart = "ensembl") # 选择人类和斑马鱼的基因数据集 human_mart <- useDataset("hsapiens_gene_ensembl", mart = ensembl) zebrafish_mart <- useDataset("drerio_gene_ensembl", mart = ensembl) # 目标人类基因列表 sample_list <- c("GREB1L","SIN3B","NCAPG2","FAM50A","PSMD12","BPTF","SLF2","SMC5", "SMC6", "TMEM260","SSBP1","TCF12", "ANLN", "TFAM", "DDX3X","REG3G") # 获取跨物种直系同源基因 orthologs <- getLDS( attributes = c("hgnc_symbol"), # 人类基因属性:HGNC符号 filters = "hgnc_symbol", # 过滤条件:HGNC符号 values = sample_list, # 目标基因列表 mart = human_mart, # 人类数据集 attributesL = c("external_gene_name", "ensembl_gene_id"), # 斑马鱼属性:基因名、Ensembl ID martL = zebrafish_mart # 斑马鱼数据集 ) # 重命名列名方便阅读 colnames(orthologs) <- c("Human_Gene", "Zebrafish_Gene_Symbol", "Zebrafish_Gene_ID") # 查看结果 print(orthologs)
说明
getLDS函数用于获取跨物种的直系同源关系- 若访问官方服务器较慢,可添加
mirror = "asia"参数切换到亚洲镜像 - 返回结果包含斑马鱼的基因符号和Ensembl ID,可根据需求调整
attributesL参数获取更多信息
内容的提问来源于stack exchange,提问作者confusedwproofs
相关产品推荐
相关产品推荐

