如何移除数组重复条目?合并同首索引元素的代码问题排查
问题:合并相同单词的条目时无法移除重复项
输入数据
new_mainArr = [ ['the', 'at', 2], ['fulton', 'np-tl', 1], ['county', 'nn-tl', 1], ['grand', 'jj-tl', 1], ['jury', 'nn-tl', 1], ['said', 'vbd', 2], ['friday', 'nr', 1], ['an', 'at', 1], ['investigation', 'nn', 1], ['of', 'in', 1], ["atlanta's", 'np$', 1], ['recent', 'jj', 1], ['primary', 'nn', 1], ['election', 'nn', 1], ['produced', 'vbd', 1], ['.', '.', 2], ['the', 'nn', 1], ['jury', 'nn', 1], ['further', 'rbr', 1], ['in', 'in', 1], ['term-end', 'nn', 1], ['presentments', 'nns', 1], ['that', 'cs', 1], ['city', 'nn-tl', 1] ]
需求说明
需要将索引0处元素相同的条目合并到同一行以移除重复项。例如'the'和'jury'各出现两次,要把它们的后续元素合并到同一行,并删除第二个'the'和'jury'条目。
现有代码及问题
编写的代码无法移除重复的条目,第二个'the'和'jury'仍然存在:
# add duplicate tags for the same word resultArr = [] temp = [] for i in range(0, len(new_mainArr)): temp = [] temp.append(new_mainArr[i][0]) temp.append(new_mainArr[i][1]) temp.append(new_mainArr[i][2]) hook = new_mainArr[i][0] for j in range(i+1, len(new_mainArr)): if(new_mainArr[j][0] == hook): temp.append(new_mainArr[j][1]) temp.append(new_mainArr[j][2]) resultArr.append(temp)
当前输出
['the', 'at', 2, 'nn', 1] ['fulton', 'np-tl', 1] ['county', 'nn-tl', 1] ['grand', 'jj-tl', 1] ['jury', 'nn-tl', 1, 'nn', 1] ['said', 'vbd', 2] ['friday', 'nr', 1] ['an', 'at', 1] ['investigation', 'nn', 1] ['of', 'in', 1] ["atlanta's", 'np$', 1] ['recent', 'jj', 1] ['primary', 'nn', 1] ['election', 'nn', 1] ['produced', 'vbd', 1] ['.', '.', 2] ['the', 'nn', 1] <= 此处不应存在 ['jury', 'nn', 1] <= 此处不应存在 ['further', 'rbr', 1] ['in', 'in', 1] ['term-end', 'nn', 1] ['presentments', 'nns', 1] ['that', 'cs', 1] ['city', 'nn-tl', 1]
解决方案
原代码的问题在于会遍历每一个元素并添加到结果中,哪怕该元素已经被前面的条目合并过。可以用集合记录已处理的单词,避免重复添加:
修改后的代码
resultArr = [] processed_words = set() # 记录已经处理过的单词 for i in range(len(new_mainArr)): word = new_mainArr[i][0] if word in processed_words: continue # 已处理过,跳过当前条目 # 初始化当前单词的合并列表 temp = [word, new_mainArr[i][1], new_mainArr[i][2]] processed_words.add(word) # 查找后续相同单词并合并元素 for j in range(i+1, len(new_mainArr)): if new_mainArr[j][0] == word: temp.append(new_mainArr[j][1]) temp.append(new_mainArr[j][2]) resultArr.append(temp) # 输出结果 for item in resultArr: print(item)
说明
processed_words集合用于跟踪已处理的单词,遍历到重复单词时直接跳过,确保每个单词只在结果中出现一次。- 仅在第一次遇到单词时创建合并列表,并收集后续所有相同单词的元素,完成合并操作。
运行结果
['the', 'at', 2, 'nn', 1] ['fulton', 'np-tl', 1] ['county', 'nn-tl', 1] ['grand', 'jj-tl', 1] ['jury', 'nn-tl', 1, 'nn', 1] ['said', 'vbd', 2] ['friday', 'nr', 1] ['an', 'at', 1] ['investigation', 'nn', 1] ['of', 'in', 1] ["atlanta's", 'np$', 1] ['recent', 'jj', 1] ['primary', 'nn', 1] ['election', 'nn', 1] ['produced', 'vbd', 1] ['.', '.', 2] ['further', 'rbr', 1] ['in', 'in', 1] ['term-end', 'nn', 1] ['presentments', 'nns', 1] ['that', 'cs', 1] ['city', 'nn-tl', 1]
内容的提问来源于stack exchange,提问作者Lyrk
相关产品推荐
相关产品推荐

