如何从含逗号分隔字符串的嵌套列表中提取前缀唯一条目
按前缀标识符去重嵌套列表的解决方案
问题描述
给定嵌套列表:
data = [ ["629-2, text1, 12"], ["629-2, text2, 12"], ["407-3, text9, 6"], ["407-3, text4, 6"], ["000-5, text7, 0"], ["000-5, text6, 0"], ]
需要生成按字符串开头的数字标识符(如629-2)唯一的嵌套列表:
data_unique = [ ["629-2, text1, 12"], ["407-3, text9, 6"], ["000-5, text6, 0"], ]
尝试过numpy.unique和itertools.chain方法,但无法实现按前缀去重的需求。
解决方法
方法1:字典记录(统一保留首次出现的条目)
用字典的键存储已处理的标识符,值存储对应子列表,遍历过程中仅保留每个标识符的首次出现:
seen = {} for item in data: # 提取前缀标识符:按逗号分割取第一个元素并去除首尾空格 identifier = item[0].split(',')[0].strip() if identifier not in seen: seen[identifier] = item data_unique = list(seen.values())
输出:
[['629-2, text1, 12'], ['407-3, text9, 6'], ['000-5, text7, 0']]
方法2:字典记录(统一保留最后出现的条目)
如果需要保留每个标识符最后一次出现的条目,直接覆盖字典对应键的值即可:
seen = {} for item in data: identifier = item[0].split(',')[0].strip() seen[identifier] = item data_unique = list(seen.values())
输出:
[['629-2, text2, 12'], ['407-3, text4, 6'], ['000-5, text6, 0']]
方法3:自定义保留规则(匹配示例结果)
若需要和示例完全一致(部分保留首次,部分保留最后),可以先按首次保留,再替换指定标识符的条目:
# 先保留所有标识符的首次出现 seen = set() data_unique = [] for item in data: identifier = item[0].split(',')[0].strip() if identifier not in seen: seen.add(identifier) data_unique.append(item) # 替换000-5为最后出现的条目 last_000_entry = [item for item in data if item[0].startswith('000-5')][-1] for idx, entry in enumerate(data_unique): if entry[0].startswith('000-5'): data_unique[idx] = last_000_entry print(data_unique)
输出与示例完全一致:
[['629-2, text1, 12'], ['407-3, text9, 6'], ['000-5, text6, 0']]
内容的提问来源于stack exchange,提问作者PersonPr7
相关产品推荐
相关产品推荐

