You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python字典统计txt文件中ENSG基因的变异数并输出CSV

问题分析与修正建议

你编写的原代码存在以下逻辑和语法错误:

  • 未定义output文件对象就调用output.write(),运行会直接报错
  • 提取ENSG编号后直接赋值为1,会覆盖之前已有的计数,完全没有实现累加逻辑
  • if 1 in d的判断逻辑完全错误,你需要判断的是提取到的ENSG编号是否在字典的键中,而非判断1是否在字典里
  • 直接把字典d赋值为整数1,会清空之前存储的所有计数数据
  • 没有按需求区分intron_variant和upstream_gene_variant,也没有输出CSV的相关逻辑
  • line.__contains__不是Python的推荐写法,更建议使用in运算符
正确实现代码
import csv

# 按需修改输入、输出文件路径
input_path = "你的基因注释文件.txt"
output_path = "基因变异计数结果.csv"

count_dict = {}

with open(input_path, 'r', encoding='utf-8') as f:
    for line in f:
        line = line.strip()
        # 跳过所有注释行和表头行
        if line.startswith('#'):
            continue
        # 按制表符分割每行字段,跳过格式异常的行
        fields = line.split('\t')
        if len(fields) < 7:
            continue
        # 提取ENSG编号(第4列,索引从0开始为3)和变异类型(第7列,索引为6)
        ensg_id = fields[3]
        variant_type = fields[6]
        # 仅统计指定的两种变异类型
        if variant_type in ('intron_variant', 'upstream_gene_variant'):
            # 累加计数
            if ensg_id in count_dict:
                count_dict[ensg_id] += 1
            else:
                count_dict[ensg_id] = 1

# 输出为CSV文件
with open(output_path, 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.writer(csv_file)
    # 写入表头
    writer.writerow(['ENSG编号', '目标变异总数量'])
    # 写入计数结果
    for ensg, count in count_dict.items():
        writer.writerow([ensg, count])
补充说明
  • 上述代码直接按字段位置提取数据,比随意切割字符串更可靠,能避免注释行或其他含ENSG的非数据行干扰计数结果
  • 如果你需要分别统计两种变异的数量,可把字典的value改为二元组存储,比如count_dict[ensg_id] = [0,0],第一个元素存intron_variant数量,第二个存upstream_gene_variant数量,对应修改累加逻辑即可

内容的提问来源于stack exchange,提问作者user15480777

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 13:36:05