You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取VCF.bgz文件报错‘Not a gzipped file’的解决方法

解决bgz格式VCF文件读取错误的方案

问题根源

你遇到的Not a gzipped file (b'TB')错误,是因为bgz是块级gzip压缩格式,和普通gzip结构存在差异,Python标准库的gzip模块无法解析这种特殊格式。

解决方法

方法1:用生物信息学专用库pysam处理(推荐)

pysam是专门处理基因组数据的Python库,原生支持bgz格式的VCF文件,还能正确解析VCF的结构:

import pysam

# 打开bgz格式的VCF输入文件
input_vcf = pysam.VariantFile("gnomad.genomes.r2.1.1.sites.2.vcf.bgz")
# 创建输出VCF文件,复用输入文件的表头
output_vcf = pysam.VariantFile("truncated.vcf", "w", header=input_vcf.header)

extracted_count = 0
for record in input_vcf:
    output_vcf.write(record)
    extracted_count += 1
    if extracted_count >= 100000:
        break

# 关闭文件
input_vcf.close()
output_vcf.close()

安装pysam:pip install pysam

方法2:用命令行工具先处理,再用Python读取

用bgzip工具(通常随samtools一起安装)先提取前10万行到普通VCF文件:

bgzip -cd gnomad.genomes.r2.1.1.sites.2.vcf.bgz | head -n 100000 > truncated.vcf

之后直接用Python读取truncated.vcf即可:

with open("truncated.vcf", "r") as f:
    for line in f:
        # 按需处理每行内容
        pass

额外检查:验证文件完整性

如果上述方法仍报错,可能是文件下载不完整。用bgzip验证文件:

bgzip -t gnomad.genomes.r2.1.1.sites.2.vcf.bgz

若输出错误提示,重新下载该文件即可。

内容的提问来源于stack exchange,提问作者Eliza Romanski

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 04:05:40