You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在mapply中使用gsub为Seurat对象指定自定义名称

问题:Seurat对象命名未按预期去除文件名后缀

问题背景

我有三组文件,需要读取后用Seurat包的CreateSeuratObject函数生成对象,希望把矩阵文件名去掉_filtered_peak_bc_matrix.h5后缀后,作为对应Seurat对象的名称。

现有代码

我编写了处理多文件的函数,但gsub未生效,生成的对象仍使用完整文件名:

read_atac <- function(sc_matrix, sc_frag, sc_meta){
  
  name <-gsub('filtered_peak_bc_matrix.h5','',sc_matrix)
  
  counts <- Read10X_h5(filename = sc_matrix)
  
  chrom_assay <- CreateChromatinAssay(
    counts = counts,
    sep = c(":", "-"),
    genome = 'hg38',
    fragments = sc_frag,
    min.cells = 10,
    min.features = 200
  )
  
  meta <- read.csv(
    file = sc_meta, 
    header = TRUE, 
    row.names = 1)
  
 assign(name, CreateSeuratObject(counts = chrom_assay, assay = "peaks", meta.data = meta)) 
  
}

mapply(read_atac, sc_matrix, sc_frag, sc_meta)

问题现象

运行后生成的Seurat对象列表键名仍是完整文件名,例如:

$`GSM8002547_Chr-Veh_R1_filtered_peak_bc_matrix.h5`
An object of class Seurat 
268284 features across 4316 samples within 1 assay 
Active assay: peaks (268284 features, 0 variable features)
 2 layers present: counts, data

我期望的对象名称是GSM8002547_Chr-Veh_R1这类格式。我的sc_matrix文件列表如下:

> sc_matrix
[1] "GSM8002547_Chr-Veh_R1_filtered_peak_bc_matrix.h5"
[2] "GSM8002548_Chr-Veh_R2_filtered_peak_bc_matrix.h5"
[3] "GSM8002549_Veh-Veh_R1_filtered_peak_bc_matrix.h5"
[4] "GSM8002550_Veh-Veh_R2_filtered_peak_bc_matrix.h5"

解决方案

问题核心是:mapply返回的列表键名由原始输入参数(即完整文件名)决定,和你用assign创建的对象名称无关——assign只是把对象放到当前环境,但不会改变mapply输出列表的结构。另外你的gsub表达式遗漏了后缀前的下划线,会导致处理后的名称末尾多一个下划线。

方案1:修改输出列表的键名

先运行函数得到结果列表,再手动替换键名:

# 运行函数获取结果(加SIMPLIFY=FALSE确保返回列表)
seurat_list <- mapply(read_atac, sc_matrix, sc_frag, sc_meta, SIMPLIFY = FALSE)
# 生成符合预期的名称
new_names <- gsub('_filtered_peak_bc_matrix.h5', '', sc_matrix)
# 替换列表的键名
names(seurat_list) <- new_names

方案2:调整函数逻辑,让输出自动继承命名

修改函数使其直接返回Seurat对象,再给输入的sc_matrix设置名称,最后用Map(即mapply的列表输出版本)运行:

# 修改函数,返回创建好的Seurat对象
read_atac <- function(sc_matrix, sc_frag, sc_meta){
  counts <- Read10X_h5(filename = sc_matrix)
  
  chrom_assay <- CreateChromatinAssay(
    counts = counts,
    sep = c(":", "-"),
    genome = 'hg38',
    fragments = sc_frag,
    min.cells = 10,
    min.features = 200
  )
  
  meta <- read.csv(
    file = sc_meta, 
    header = TRUE, 
    row.names = 1)
  
  CreateSeuratObject(counts = chrom_assay, assay = "peaks", meta.data = meta)
}

# 给sc_matrix设置期望的名称
names(sc_matrix) <- gsub('_filtered_peak_bc_matrix.h5', '', sc_matrix)
# 用Map运行,输出列表会自动继承设置好的名称
seurat_list <- Map(read_atac, sc_matrix, sc_frag, sc_meta)

内容的提问来源于stack exchange,提问作者Kousalya Devi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 05:12:47