Python字符串字符占比代码问题:需输出最终结果而非逐行占比
DNA序列GC占比统计代码修正
我要统计DNA序列中C、G字符的出现占比,但现有Python代码会逐行输出百分比,不符合需求——我需要每个Ros_xxxx条目对应输出最终的占比结果。
原代码
from os import cpu_count import re with open('ros_gc_948_1_dataset.txt', 'r') as fp: # read all lines in a list lines = fp.readlines() c_count=0 g_count=0 a_count=0 t_count=0 for line in lines: # check if string present on a current line if re.findall(r"\bRos_[0-9][0-9][0-9][0-9]" , line): line.write() line.replace(">","",) print(line) c_count=0 g_count=0 a_count=0 t_count=0 else: c_count+=line.count("C") g_count+=line.count("G") a_count+=line.count("A") t_count+=line.count("T") tocag=c_count+g_count toac=c_count + g_count + a_count + t_count if toac !=0: avrg=float(tocag/toac * 100) print(avrg)
当前输出
>Ros_7657 53.333333333333336 43.333333333333336 47.22222222222222 48.333333333333336 49.0 49.72222222222222 49.76190476190476 48.541666666666664 48.333333333333336 48.833333333333336 49.24242424242424 49.30555555555556 48.717948717948715 49.25925925925926 >Ros_3487 53.333333333333336 47.5 46.666666666666664 48.333333333333336 47.0 46.94444444444444 48.095238095238095 47.291666666666664 46.111111111111114 47.333333333333336 46.96969696969697 46.52777777777778 46.666666666666664 46.785714285714285 47.368421052631575
期望输出
Ros_7657 49.25925925925926 Ros_3487 47.368421052631575
输入文件示例
>Ros_2115 GAGGCAATGGTTATCAACCCCTGATTTACGAATGACCTAACAACTCCTTAGAATTTAATC GTTATGTGAATTAAGCAACGCTCGCGAATTGCTATGTTAATTCGCACTGTAAGGTGTCGA ACGAAATCCACTGTTCCTTTTCTAATTTCTTTCA
修正后的代码
import re with open('ros_gc_948_1_dataset.txt', 'r') as fp: lines = fp.readlines() c_count = 0 g_count = 0 a_count = 0 t_count = 0 current_id = "" for line in lines: line = line.strip() # 去除换行符和首尾空格 if not line: continue # 跳过空行 # 匹配Ros_xxxx格式的条目ID match = re.search(r"Ros_\d{4}", line) if match: # 若已有未输出的条目,先输出其最终占比 if current_id: total = c_count + g_count + a_count + t_count if total != 0: gc_ratio = (c_count + g_count) / total * 100 print(current_id) print(gc_ratio) # 更新当前条目ID,重置计数器 current_id = match.group() c_count = 0 g_count = 0 a_count = 0 t_count = 0 else: # 累计当前行的碱基数量 c_count += line.count("C") g_count += line.count("G") a_count += line.count("A") t_count += line.count("T") # 处理最后一个未输出的条目 if current_id: total = c_count + g_count + a_count + t_count if total != 0: gc_ratio = (c_count + g_count) / total * 100 print(current_id) print(gc_ratio)
修正说明
- 移除了无用的
cpu_count导入,原代码未使用该模块 - 新增
current_id变量跟踪当前处理的条目ID,解决条目切换时的输出逻辑问题 - 遇到新条目时先输出上一个条目的最终占比,避免逐行输出中间结果
- 循环结束后处理最后一个条目,防止遗漏输出
- 修复了原代码中
line.write()的错误调用,改用正则提取条目ID - 加入空行跳过逻辑,避免空行干扰统计结果
内容的提问来源于stack exchange,提问作者abolfazlmalekahmadi
相关产品推荐
相关产品推荐

