Python遍历嵌套列表子元素配对提取序列时索引错误,如何优化实现?
错误原因
- 首要原因是
populated_indices的长度和pca_df的行数不匹配:当循环中index的值超过populated_indices的最大下标时,取populated_indices[index]就会抛出IndexError - 次要原因是双指针的遍历写法冗余,容易出现边界逻辑错误
优化实现方案
基础实现(无需引入额外库)
直接利用列表步长切片配对,逻辑清晰不易出错:
lst_separated = [['1605',1607],['230320','230400','25020','26232'],['230320','230400','25020','26232','302320','312320'],...] # 提前做长度校验,从根源避免越界 assert len(populated_indices) == len(pca_df), "populated_indices长度和pca_df行数不匹配" populated_lists = [lst_separated[i] for i in populated_indices] outcome = [] for seqName, sublist in zip(pca_df['seqName'], populated_lists): # 两两配对:取奇数位元素和偶数位元素一一对应,长度为奇数时自动丢弃最后一个不匹配的元素 for pos1, pos2 in zip(sublist[::2], sublist[1::2]): start = int(pos1) - 1 end = int(pos2) - 1 var = str(record_dict_2[seqName].seq)[start:end] outcome.append(var)
itertools实现(适合大列表场景,内存效率更高)
用迭代器实现配对,避免生成临时切片,超大子列表场景下性能优势更明显:
from itertools import islice lst_separated = [['1605',1607],['230320','230400','25020','26232'],['230320','230400','25020','26232','302320','312320'],...] assert len(populated_indices) == len(pca_df), "populated_indices长度和pca_df行数不匹配" populated_lists = [lst_separated[i] for i in populated_indices] outcome = [] for seqName, sublist in zip(pca_df['seqName'], populated_lists): iter_s = iter(sublist) # 迭代器每次调用都会自动移动指针,两次取同一个迭代器的值自然实现两两配对 for pos1, pos2 in zip(iter_s, iter_s): start = int(pos1) - 1 end = int(pos2) - 1 var = str(record_dict_2[seqName].seq)[start:end] outcome.append(var)
内容的提问来源于stack exchange,提问作者user17153595
相关产品推荐
相关产品推荐

