如何编写代码识别线性化RNA结构中茎环的起止位置?
问题分析与代码修正
你的代码当前会捕获每个">"块的第一个">",但我们需要的是每个">"块的最后一个">"。要解决这个问题,我们需要调整逻辑,等待遍历到">"块的末尾时再记录结束索引。
修正后的代码
RNA_SS = "..<<<<...>>>>..<<..>>" RNA_seq = "..AAUGCCCCAUU..CCAAGG" def id_stem_loops(consensus): stem_loops = [] skip_sym = {".", ",", "-", "~", "_"} # 用集合提升查找效率 start_idx = None total_length = len(consensus) for idx, sym in enumerate(consensus): if sym in skip_sym: continue # 识别连续"<"块的起始位置 if sym == "<": # 当前是"<"且前一个字符不是"<",说明是新的"<"块开始 if idx == 0 or consensus[idx-1] != "<": start_idx = idx # 识别连续">"块的结束位置 elif sym == ">": # 当前是">"且下一个字符不是">"(或已是最后一个字符),说明是">"块的结尾 if idx == total_length - 1 or consensus[idx+1] != ">": if start_idx is not None: stem_loops.append((start_idx, idx)) start_idx = None # 配对完成后重置起始索引 print(stem_loops) id_stem_loops(RNA_SS)
关键逻辑调整
- 跟踪"<"块的起始:当遇到连续"<"的第一个字符时,记录其索引作为茎环的起始位置。
- 等待">"块的结束:只有当遍历到连续">"的最后一个字符时,才将之前记录的起始索引与当前索引配对,加入结果列表。
- 重置起始索引:每完成一个茎环的配对后,重置起始索引,确保下一个"<"块能正确被捕获。
运行修正后的代码,输出将是期望的 [(2, 12), (15, 20)]。
内容的提问来源于stack exchange,提问作者Kate
相关产品推荐
相关产品推荐

