You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在保留顺序的前提下移除列表中特定格式的重复字符串?

特定格式元素的列表去重(保留原顺序)

给定如下Python字符串列表:

list_ex = ['I', 'went', 'to', 'the', 'big', 'conference', ',', 'I', 'presented', 'myself', 'there', '.', 'After', 'the', '<word>conference</word>', '<word>conference</word>', ',', 'I', 'took', 'a', 'taxi', 'to', 'go', 'to', 'the', '<word>hotel</word>', '<word>hotel</word>', '.']

需要移除其中格式为<word>keyword</word>的重复字符串,得到目标列表:

new_list_ex = ['I', 'went', 'to', 'the', 'big', 'conference', ',', 'I', 'presented', 'myself', 'there', '.', 'After', 'the', '<word>conference</word>', ',', 'I', 'took', 'a', 'taxi', 'to', 'go', 'to', 'the', '<word>hotel</word>', '.']

已知常规列表去重方法,如何针对这类特定元素去重并保留原列表顺序?


解决方案

核心思路:遍历原列表时,仅对符合<word>...</word>格式的元素做去重判断,其他元素直接保留;用集合记录已保留的特定元素,避免重复添加。

方法1:基础遍历实现

list_ex = ['I', 'went', 'to', 'the', 'big', 'conference', ',', 'I', 'presented', 'myself', 'there', '.', 'After', 'the', '<word>conference</word>', '<word>conference</word>', ',', 'I', 'took', 'a', 'taxi', 'to', 'go', 'to', 'the', '<word>hotel</word>', '<word>hotel</word>', '.']

new_list = []
seen_special = set()

for item in list_ex:
    # 判断当前元素是否为目标格式
    if item.startswith('<word>') and item.endswith('</word>'):
        if item not in seen_special:
            new_list.append(item)
            seen_special.add(item)
    else:
        # 非目标格式元素直接加入结果列表
        new_list.append(item)

print(new_list)

方法2:生成器表达式简化代码

如果想让代码更紧凑,可以结合生成器和辅助函数实现:

list_ex = ['I', 'went', 'to', 'the', 'big', 'conference', ',', 'I', 'presented', 'myself', 'there', '.', 'After', 'the', '<word>conference</word>', '<word>conference</word>', ',', 'I', 'took', 'a', 'taxi', 'to', 'go', 'to', 'the', '<word>hotel</word>', '<word>hotel</word>', '.']

seen_special = set()

def should_keep(item):
    if item.startswith('<word>') and item.endswith('</word>'):
        if item not in seen_special:
            seen_special.add(item)
            return True
        return False
    return True

new_list = [item for item in list_ex if should_keep(item)]

print(new_list)

说明

  • 两种方法都只对<word>...</word>格式的元素去重,其他元素无论是否重复都会原样保留,完全匹配需求。
  • 用集合seen_special记录已添加的特定元素,查询效率为O(1),整体时间复杂度是O(n),处理大列表也能保证高效。

内容的提问来源于stack exchange,提问作者Erwin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 01:35:18