You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python字符串字符占比代码问题:需输出最终结果而非逐行占比

DNA序列GC占比统计代码修正

我要统计DNA序列中C、G字符的出现占比,但现有Python代码会逐行输出百分比,不符合需求——我需要每个Ros_xxxx条目对应输出最终的占比结果。

原代码

from os import cpu_count
import re

with open('ros_gc_948_1_dataset.txt', 'r') as fp:
    # read all lines in a list
    lines = fp.readlines()
    c_count=0
    g_count=0
    a_count=0
    t_count=0
    for line in lines:

        # check if string present on a current line
        if  re.findall(r"\bRos_[0-9][0-9][0-9][0-9]" , line):
            line.write()
            line.replace(">","",)
            print(line)

            c_count=0
            g_count=0
            a_count=0
            t_count=0
        else:
           c_count+=line.count("C")
           g_count+=line.count("G")
           a_count+=line.count("A")
           t_count+=line.count("T")
           
        tocag=c_count+g_count
        toac=c_count + g_count + a_count + t_count
        if toac !=0:
         avrg=float(tocag/toac * 100)
         print(avrg)

当前输出

>Ros_7657

53.333333333333336
43.333333333333336
47.22222222222222
48.333333333333336
49.0
49.72222222222222
49.76190476190476
48.541666666666664
48.333333333333336
48.833333333333336
49.24242424242424
49.30555555555556
48.717948717948715
49.25925925925926
>Ros_3487

53.333333333333336
47.5
46.666666666666664
48.333333333333336
47.0
46.94444444444444
48.095238095238095
47.291666666666664
46.111111111111114
47.333333333333336
46.96969696969697
46.52777777777778
46.666666666666664
46.785714285714285
47.368421052631575

期望输出

Ros_7657
49.25925925925926
Ros_3487
47.368421052631575

输入文件示例

>Ros_2115
GAGGCAATGGTTATCAACCCCTGATTTACGAATGACCTAACAACTCCTTAGAATTTAATC
GTTATGTGAATTAAGCAACGCTCGCGAATTGCTATGTTAATTCGCACTGTAAGGTGTCGA
ACGAAATCCACTGTTCCTTTTCTAATTTCTTTCA

修正后的代码

import re

with open('ros_gc_948_1_dataset.txt', 'r') as fp:
    lines = fp.readlines()
    c_count = 0
    g_count = 0
    a_count = 0
    t_count = 0
    current_id = ""
    
    for line in lines:
        line = line.strip()  # 去除换行符和首尾空格
        if not line:
            continue  # 跳过空行
        
        # 匹配Ros_xxxx格式的条目ID
        match = re.search(r"Ros_\d{4}", line)
        if match:
            # 若已有未输出的条目,先输出其最终占比
            if current_id:
                total = c_count + g_count + a_count + t_count
                if total != 0:
                    gc_ratio = (c_count + g_count) / total * 100
                    print(current_id)
                    print(gc_ratio)
            
            # 更新当前条目ID,重置计数器
            current_id = match.group()
            c_count = 0
            g_count = 0
            a_count = 0
            t_count = 0
        else:
            # 累计当前行的碱基数量
            c_count += line.count("C")
            g_count += line.count("G")
            a_count += line.count("A")
            t_count += line.count("T")
    
    # 处理最后一个未输出的条目
    if current_id:
        total = c_count + g_count + a_count + t_count
        if total != 0:
            gc_ratio = (c_count + g_count) / total * 100
            print(current_id)
            print(gc_ratio)

修正说明

  1. 移除了无用的cpu_count导入,原代码未使用该模块
  2. 新增current_id变量跟踪当前处理的条目ID,解决条目切换时的输出逻辑问题
  3. 遇到新条目时先输出上一个条目的最终占比,避免逐行输出中间结果
  4. 循环结束后处理最后一个条目,防止遗漏输出
  5. 修复了原代码中line.write()的错误调用,改用正则提取条目ID
  6. 加入空行跳过逻辑,避免空行干扰统计结果

内容的提问来源于stack exchange,提问作者abolfazlmalekahmadi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 14:01:06