htseq-count无法生成预期读数计数,GFF文件属性缺失报错求助
问题描述
我有如下GFF文件:
caffold1 GeneWise mRNA 227302 283623 80.88 - . ID=Mnat_00001;evid_id=ENST00000360911;Shift=0; scaffold1 GeneWise CDS 227302 227498 . - 2 Parent=Mnat_00001; scaffold1 GeneWise CDS 230150 230298 . - 1 Parent=Mnat_00001; scaffold1 GeneWise CDS 233743 234426 . - 1 Parent=Mnat_00001; scaffold1 GeneWise CDS 236092 236835 . - 1 Parent=Mnat_00001; scaffold1 GeneWise CDS 238558 238807 . - 2 Parent=Mnat_00001; scaffold1 GeneWise CDS 240781 240970 . - 0 Parent=Mnat_00001; scaffold1 GeneWise CDS 241779 241912 . - 2 Parent=Mnat_00001; scaffold1 GeneWise CDS 249825 250005 . - 0 Parent=Mnat_00001; scaffold1 GeneWise CDS 273368 273452 . - 1 Parent=Mnat_00001; scaffold1 GeneWise CDS 283460 283623 . - 0 Parent=Mnat_00001; scaffold1 Cuff mRNA 316222 341723 1000 + . ID=Mnat_00002;evid_id=CCG000108.1;source_id=CUFF.1168.1; scaffold1 Cuff CDS 316222 316368 1000 + 0 Parent=Mnat_00002; scaffold1 Cuff CDS 22468 322630 1000 + 0 Parent=Mnat_00002; scaffold1 Cuff CDS 325120 325274 1000 + 2 Parent=Mnat_00002; scaffold1 Cuff CDS 329797 329922 1000 + 0 Parent=Mnat_00002;
执行以下bash命令进行RNA-seq读数计数:
for i in *sorted.bam; do htseq-count $i Mnat_gene_v1.2.gff -f bam -i -mRNA -t CDS -m union -r name --stranded=no > ${i}_count.txt; done
出现报错:
Error processing GFF file (line 2 of file Mnat_gene_v1.2.gff): Feature Mnat_00001 does not contain a 'mRNA' attribute [Exception type: ValueError, raised in features.py:387]
错误原因
核心问题是htseq-count的参数配置错误:
-i参数用于指定GFF文件中作为特征唯一标识的属性名,你写成了-i -mRNA,这会让程序错误地寻找名为mRNA的属性,但你的GFF里特征的标识属性是ID(比如mRNA行的ID=Mnat_00001,CDS行通过Parent=Mnat_00001关联到对应mRNA)。- 错误的
-i参数导致程序无法正确识别CDS对应的父特征,进而触发属性缺失的报错。
解决方法
修正htseq-count的参数配置:
- 将
-i -mRNA改为-i ID,明确告诉程序使用GFF中的ID字段作为特征标识; - 若目标是统计mRNA转录本的读数,将
-t CDS改为-t mRNA(因为CDS是mRNA的子特征,通常转录本或基因层级的计数更常用);如果确实需要统计CDS的读数,保留-t CDS即可。
统计mRNA转录本的修正命令
for i in *sorted.bam; do htseq-count $i Mnat_gene_v1.2.gff -f bam -i ID -t mRNA -m union -r name --stranded=no > ${i}_count.txt; done
统计CDS的修正命令
for i in *sorted.bam; do htseq-count $i Mnat_gene_v1.2.gff -f bam -i ID -t CDS -m union -r name --stranded=no > ${i}_count.txt; done
补充注意事项
- 确认你的BAM文件确实是按名称排序的(对应参数
-r name),如果是按位置排序的,需改为-r pos; --stranded=no适用于无链特异性的建库方案,请根据你的实验设计确认参数是否匹配。
内容的提问来源于stack exchange,提问作者Allan Okwaro
相关产品推荐
相关产品推荐

