能否将合并重复歌曲的嵌套循环改写为嵌套列表推导式?
合并CSV中重复歌曲:列表推导式优化方案
我正在处理一个包含大量歌曲信息的CSV文件,需要合并其中的重复歌曲:
masterList是一个由列表组成的元组,每个列表代表一首歌曲。songsMatch()函数接收两首歌曲作为输入,根据二者是否匹配输出False或True。
原本用列表推导式只能生成两两配对的子列表,无法将所有匹配歌曲归到同一子列表;后来写的循环代码基本可行,但转成嵌套列表推导式时遇到两个问题:
- 输出的
matches包含大量空列表 - 子列表仅包含匹配的后续歌曲,缺失
masterList[index1]对应的首项
原尝试的列表推导式:
matches=[[masterList[index2] for index2 in range(index1+1, len(masterList)) if songsMatch(masterList[index1], masterList[index2])] for index1 in range(1, len(masterList)-2)]
一、修复列表推导式的核心问题
要同时解决空列表和缺失首项的问题,可先构造包含当前歌曲的基础列表,再拼接匹配项,最后过滤掉无匹配的单元素列表:
matches = [ [masterList[index1]] + [masterList[index2] for index2 in range(index1+1, len(masterList)) if songsMatch(masterList[index1], masterList[index2])] for index1 in range(1, len(masterList)-2) if any(songsMatch(masterList[index1], masterList[index2]) for index2 in range(index1+1, len(masterList))) ]
关键说明:
[masterList[index1]] + [...]:将当前歌曲作为子列表首项,再拼接所有匹配的后续歌曲if any(...):过滤掉没有任何匹配项的情况,避免生成仅含单个元素的无效子列表
二、解决原逻辑的重复分组问题(可选优化)
原循环代码存在缺陷:若某首歌曲重复3次(如A、B、C互相匹配),会生成[A,B,C]和[B,C]两个重复分组。如果想要完全去重,可通过标记已处理歌曲避免重复遍历,虽然这不是纯列表推导式,但可读性和实用性更强:
processed = set() matches = [] for index1 in range(1, len(masterList)-2): if index1 in processed: continue group = [masterList[index1]] for index2 in range(index1+1, len(masterList)): if songsMatch(masterList[index1], masterList[index2]): group.append(masterList[index2]) processed.add(index2) if len(group) > 1: matches.append(group)
复杂的分组逻辑用循环维护更直观,没必要强行嵌套列表推导式。
内容的提问来源于stack exchange,提问作者Smiley Tiger
相关产品推荐
相关产品推荐

