You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除Python列表中特定字符串的相邻重复项

解决Python列表中特定字符串的相邻重复项移除问题

问题场景

现有示例列表:

list_ex = ['I', 'went', 'to', 'the', 'big', 'conference', ',', 'I', 'presented', 'myself', 'there', '.', 'After', 'the', '<word>conference</word>', '<word>conference</word>', ',', 'I', 'took', 'a', 'taxi', 'to', 'go', 'to', 'the', '<word>hotel</word>', '<word>hotel</word>', '.', 'Tomorrow', 'I', 'will', 'go', 'to', '<word>conference</word>', 'again', '.']

尝试的错误代码:

new_list_ex = []
for item in list_ex:
    if item.startswith('<word>'):
        if item in new_list_ex and (item == list_ex[list_ex.index(item)+1]):
            continue
    new_list_ex.append(item)

得到的错误输出:

['I', 'went', 'to', 'the', 'big', 'conference', ',', 'I', 'presented', 'myself', 'there', '.', 'After', 'the', '<word>conference</word>', ',', 'I', 'took', 'a', 'taxi', 'to', 'go', 'to', 'the', '<word>hotel</word>', '.', 'Tomorrow', 'I', 'will', 'go', 'to', 'again', '.']

期望输出:

['I', 'went', 'to', 'the', 'big', 'conference', ',', 'I', 'presented', 'myself', 'there', '.', 'After', 'the', '<word>conference</word>', ',', 'I', 'took', 'a', 'taxi', 'to', 'go', 'to', 'the', '<word>hotel</word>', '.', 'Tomorrow', 'I', 'will', 'go', 'to', '<word>conference</word>', 'again', '.']

问题根源

你指出的list_ex[list_ex.index(item)+1]逻辑确实存在问题:

  • list_ex.index(item)返回的是列表中第一个匹配该元素的索引,而非当前循环到的元素位置。比如最后一个<word>conference</word>,index()会返回它第一次出现的索引(14),进而取索引15的元素(同样是<word>conference</word>),导致代码误判当前元素重复,从而跳过添加,这就是最后那个标签丢失的原因。
  • item in new_list_ex的判断完全多余,我们要移除的是相邻的重复项,只需关注当前元素和前一个被添加到新列表的元素是否相同即可。

修正方案

方案1:跟踪新列表的最后一个元素

遍历列表时,仅当当前元素不是目标类型(<word>开头),或者当前元素和新列表最后一个元素不同时,才添加到新列表:

new_list_ex = []
for item in list_ex:
    # 如果是目标类型,且新列表不为空,且最后一个元素和当前元素相同,则跳过
    if item.startswith('<word>') and new_list_ex and new_list_ex[-1] == item:
        continue
    new_list_ex.append(item)

方案2:带索引遍历列表

通过enumerate获取当前元素的索引,直接比较当前元素和下一个元素(注意边界判断):

new_list_ex = []
for idx, item in enumerate(list_ex):
    # 如果是目标类型,且不是最后一个元素,且下一个元素和当前元素相同,则跳过
    if item.startswith('<word>') and idx < len(list_ex)-1 and list_ex[idx+1] == item:
        continue
    new_list_ex.append(item)

两种方案都能得到期望输出,其中方案1更简洁高效,完全贴合“移除相邻重复项”的需求。

内容的提问来源于stack exchange,提问作者Erwin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 03:10:14