You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行GSEApy触发AssertionError: assert len(dat) >1问题求助

GSEApy运行GSEA时触发AssertionError: assert len(dat) >1的解决方法

问题背景

使用GSEApy导入自定义.cls文件执行GSEA分析时,触发断言错误AssertionError: assert len(dat) >1,但已确认输入的DataFrame行数大于1,无法定位问题根源。


相关代码与报错信息

生成.cls文件的代码

groups = subtype 
with open("./kirp_deg_gene_exp.cls", "w") as cl:
    line = f"{len(groups)}\t2\t1\n#\tKIRP\tnormal\n"
    cl.write(line)
    cl.write("\t".join(groups) + "\t")

gene_sets = ["KEGG_2021_Human", "Reactome_2022", "BioCarta_2016", "WikiPathway_2021_Human", "Panther_2016"]

运行GSEA的代码

gs_res = gp.gsea(data=rcc_normal_deg, # or data='./P53_resampling_data.txt'
                 gene_sets=gene_sets, # or enrichr library names
                 cls= "./kirp_deg_gene_exp.cls", # cls=class_vector
                 # set permutation_type to phenotype if samples >=15
                 permutation_type='phenotype',
                 permutation_num=1000, # reduce number to speed up test
                 outdir=None,  # do not write output to disk
                 method='signal_to_noise',
                 threads=4, seed= 7)

报错回溯

---------------------------------------------------------------------------
AssertionError                            Traceback (most recent call last)
/tmp/ipykernel_27/3071075505.py in <module>
      9                  outdir=None,  # do not write output to disk
     10                  method='signal_to_noise',
---> 11                  threads=4, seed= 7)

/opt/conda/lib/python3.7/site-packages/gseapy/__init__.py in gsea(data, gene_sets, cls, outdir, min_size, max_size, permutation_num, weighted_score_type, permutation_type, method, ascending, threads, figsize, format, graph_num, no_plot, seed, verbose, *arg, **kwarg)
    147         verbose,
    148     )
---> 149     gs.run()
    150 
    151     return gs

/opt/conda/lib/python3.7/site-packages/gseapy/gsea.py in run(self)
    258         self.cls_dict = cls_dict
    259         # data frame must have length > 1
---> 260         assert len(dat) > 1
    261         # filtering out gene sets and build gene sets dictionary
    262         gmt = self.load_gmt(gene_list=dat.index.values, gmt=self.gene_sets)

AssertionError: 

数据示例

{'TCGA-Y8-A8RY-01A': {'ZNF771': '131.6026',
  'ZNF775': '328.4314',
  'ZNF789': '568.6275',
  'subtype': 'KIRP'},
 'TCGA-Y8-A8RY-11A': {'ZNF771': '61.4095',
  'ZNF775': '107.5387',
  'ZNF789': '63.4277',
  'subtype': 'normal'},
 'TCGA-Y8-A8RZ-01A': {'ZNF771': '171.2236',
  'ZNF775': '96.9937',
  'ZNF789': '93.5296',
  'subtype': 'KIRP'},
 'TCGA-Y8-A8S0-01A': {'ZNF771': '107.6399',
  'ZNF775': '198.0431',
  'ZNF789': '249.7065',
  'subtype': 'KIRP'}}

问题排查与解决方案

1. 表达值类型错误(核心原因)

从数据示例可见,所有表达值均为字符串格式(如'131.6026'),而非数值类型。GSEApy会自动过滤无法转换为数值的基因,最终导致剩余有效基因数≤1,触发断言错误。

解决方法:将DataFrame中的表达值转换为数值类型:

# 排除非表达列(如subtype),将其余列转为float
expr_cols = rcc_normal_deg.columns.difference(['subtype'])
rcc_normal_deg[expr_cols] = rcc_normal_deg[expr_cols].astype(float)

2. .cls文件格式错误

生成.cls文件时,最后一行多写了一个额外的\t,可能导致分组数量与样本数不匹配,间接引发数据处理异常。

解决方法:修正.cls文件写入代码,移除末尾多余的\t:

groups = subtype 
with open("./kirp_deg_gene_exp.cls", "w") as cl:
    line = f"{len(groups)}\t2\t1\n#\tKIRP\tnormal\n"
    cl.write(line)
    cl.write("\t".join(groups))  # 移除末尾多余的\t

3. 数据行列方向错误

GSEA要求输入的表达矩阵必须是基因作为行索引,样本作为列。若DataFrame是样本作为行、基因作为列,会导致GSEApy识别的有效基因数为0或1。

解决方法:若行列方向错误,转置DataFrame并排除非表达列:

# 确保行是基因,列是样本,同时移除subtype列
rcc_normal_deg = rcc_normal_deg.drop('subtype', axis=1).T

验证步骤

  1. 转换数据类型后,执行print(rcc_normal_deg.dtypes),确认所有表达列均为float或int类型。
  2. 打开生成的.cls文件,检查最后一行的分组数量与样本数完全一致,无多余制表符。
  3. 执行print(rcc_normal_deg.head()),确认DataFrame结构为:行是基因名称,列是样本ID。

内容的提问来源于stack exchange,提问作者melolilili

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 02:40:39