运行GSEApy触发AssertionError: assert len(dat) >1问题求助
GSEApy运行GSEA时触发
AssertionError: assert len(dat) >1的解决方法 问题背景
使用GSEApy导入自定义.cls文件执行GSEA分析时,触发断言错误AssertionError: assert len(dat) >1,但已确认输入的DataFrame行数大于1,无法定位问题根源。
相关代码与报错信息
生成.cls文件的代码
groups = subtype with open("./kirp_deg_gene_exp.cls", "w") as cl: line = f"{len(groups)}\t2\t1\n#\tKIRP\tnormal\n" cl.write(line) cl.write("\t".join(groups) + "\t") gene_sets = ["KEGG_2021_Human", "Reactome_2022", "BioCarta_2016", "WikiPathway_2021_Human", "Panther_2016"]
运行GSEA的代码
gs_res = gp.gsea(data=rcc_normal_deg, # or data='./P53_resampling_data.txt' gene_sets=gene_sets, # or enrichr library names cls= "./kirp_deg_gene_exp.cls", # cls=class_vector # set permutation_type to phenotype if samples >=15 permutation_type='phenotype', permutation_num=1000, # reduce number to speed up test outdir=None, # do not write output to disk method='signal_to_noise', threads=4, seed= 7)
报错回溯
--------------------------------------------------------------------------- AssertionError Traceback (most recent call last) /tmp/ipykernel_27/3071075505.py in <module> 9 outdir=None, # do not write output to disk 10 method='signal_to_noise', ---> 11 threads=4, seed= 7) /opt/conda/lib/python3.7/site-packages/gseapy/__init__.py in gsea(data, gene_sets, cls, outdir, min_size, max_size, permutation_num, weighted_score_type, permutation_type, method, ascending, threads, figsize, format, graph_num, no_plot, seed, verbose, *arg, **kwarg) 147 verbose, 148 ) ---> 149 gs.run() 150 151 return gs /opt/conda/lib/python3.7/site-packages/gseapy/gsea.py in run(self) 258 self.cls_dict = cls_dict 259 # data frame must have length > 1 ---> 260 assert len(dat) > 1 261 # filtering out gene sets and build gene sets dictionary 262 gmt = self.load_gmt(gene_list=dat.index.values, gmt=self.gene_sets) AssertionError:
数据示例
{'TCGA-Y8-A8RY-01A': {'ZNF771': '131.6026', 'ZNF775': '328.4314', 'ZNF789': '568.6275', 'subtype': 'KIRP'}, 'TCGA-Y8-A8RY-11A': {'ZNF771': '61.4095', 'ZNF775': '107.5387', 'ZNF789': '63.4277', 'subtype': 'normal'}, 'TCGA-Y8-A8RZ-01A': {'ZNF771': '171.2236', 'ZNF775': '96.9937', 'ZNF789': '93.5296', 'subtype': 'KIRP'}, 'TCGA-Y8-A8S0-01A': {'ZNF771': '107.6399', 'ZNF775': '198.0431', 'ZNF789': '249.7065', 'subtype': 'KIRP'}}
问题排查与解决方案
1. 表达值类型错误(核心原因)
从数据示例可见,所有表达值均为字符串格式(如'131.6026'),而非数值类型。GSEApy会自动过滤无法转换为数值的基因,最终导致剩余有效基因数≤1,触发断言错误。
解决方法:将DataFrame中的表达值转换为数值类型:
# 排除非表达列(如subtype),将其余列转为float expr_cols = rcc_normal_deg.columns.difference(['subtype']) rcc_normal_deg[expr_cols] = rcc_normal_deg[expr_cols].astype(float)
2. .cls文件格式错误
生成.cls文件时,最后一行多写了一个额外的\t,可能导致分组数量与样本数不匹配,间接引发数据处理异常。
解决方法:修正.cls文件写入代码,移除末尾多余的\t:
groups = subtype with open("./kirp_deg_gene_exp.cls", "w") as cl: line = f"{len(groups)}\t2\t1\n#\tKIRP\tnormal\n" cl.write(line) cl.write("\t".join(groups)) # 移除末尾多余的\t
3. 数据行列方向错误
GSEA要求输入的表达矩阵必须是基因作为行索引,样本作为列。若DataFrame是样本作为行、基因作为列,会导致GSEApy识别的有效基因数为0或1。
解决方法:若行列方向错误,转置DataFrame并排除非表达列:
# 确保行是基因,列是样本,同时移除subtype列 rcc_normal_deg = rcc_normal_deg.drop('subtype', axis=1).T
验证步骤
- 转换数据类型后,执行
print(rcc_normal_deg.dtypes),确认所有表达列均为float或int类型。 - 打开生成的.cls文件,检查最后一行的分组数量与样本数完全一致,无多余制表符。
- 执行
print(rcc_normal_deg.head()),确认DataFrame结构为:行是基因名称,列是样本ID。
内容的提问来源于stack exchange,提问作者melolilili
相关产品推荐
相关产品推荐

