Python列表元素重复记录 如何实现概念同义词打印结果去重
同义词提取功能输出重复问题解决
问题描述
开发从数据集/样本中提取同义词功能时,使用现有代码打印概念与关联同义词会重复输出两次,可读性差,需要实现仅打印唯一结果。
现有问题代码
for rt in self.raw_tokens: concept = None if rt.startswith('<'): # 如果是概念(实体) if 'value' in ElementTree.fromstring(rt).attrib: string_token = ElementTree.fromstring(rt).text concept = ElementTree.fromstring(rt).tag # 如果是同义词,直接取文本并移除样本中的标签 else: string_token = ElementTree.fromstring(rt).text #print('Token: [' + rt + ']') print("Concepts & Synonyms are present in the sample :<%s> = %s" %(ElementTree.fromstring(rt).tag,string_token)) #print('CONCEPT: ' + concept) else: string_token = rt toks.append((concept, string_token))
输出对比
预期输出
Concepts & Synonyms are present in the sample :<<I_LOVE>I like</I_LOVE>> = I like Concepts & Synonyms are present in the sample :<<I_WANTS>NEED</I_WANTS>> = NEED Concepts & Synonyms are present in the sample :<<NEW_YORK>NEW YORK</NEW_YORK>> = NEW YORK Concepts & Synonyms are present in the sample :<<I_WANTS>wish</I_WANTS>> = wish Concepts & Synonyms are present in the sample :<<NEW_YORK>BIG APPLE</NEW_YORK>> = BIG APPLE
实际输出(存在重复)
Concepts & Synonyms are present in the sample :<<I_LOVE>I like</I_LOVE>> = I like Concepts & Synonyms are present in the sample :<<I_WANTS>NEED</I_WANTS>> = NEED Concepts & Synonyms are present in the sample :<<NEW_YORK>NEW YORK</NEW_YORK>> = NEW YORK Concepts & Synonyms are present in the sample :<<I_WANTS>wish</I_WANTS>> = wish Concepts & Synonyms are present in the sample :<<NEW_YORK>BIG APPLE</NEW_YORK>> = BIG APPLE Concepts & Synonyms are present in the sample :<<I_LOVE>I like</I_LOVE>> = I like Concepts & Synonyms are present in the sample :<<I_WANTS>NEED</I_WANTS>> = NEED Concepts & Synonyms are present in the sample :<<NEW_YORK>NEW YORK</NEW_YORK>> = NEW YORK Concepts & Synonyms are present in the sample :<<I_WANTS>wish</I_WANTS>> = wish Concepts & Synonyms are present in the sample :<<NEW_YORK>BIG APPLE</NEW_YORK>> = BIG APPLE
解决方案
输出重复的核心原因是self.raw_tokens本身存在重复条目,或是该段循环代码被重复调用,两种修复方案如下:
方案1:新增已打印集合去重(对原有代码改动最小)
在循环外初始化空集合记录已打印的内容,打印前先判断是否已存在,不存在才打印同时加入集合,同时优化重复解析XML的冗余逻辑:
# 初始化去重集合,放在循环外部 printed_entries = set() for rt in self.raw_tokens: concept = None if rt.startswith('<'): # 提前解析XML避免重复调用 elem = ElementTree.fromstring(rt) if 'value' in elem.attrib: string_token = elem.text concept = elem.tag else: string_token = elem.text tag = elem.tag # 用概念+同义词的组合作为唯一判断键 entry_key = (tag, string_token) if entry_key not in printed_entries: printed_entries.add(entry_key) print(f"Concepts & Synonyms are present in the sample :<{tag}> = {string_token}") else: string_token = rt toks.append((concept, string_token))
方案2:提前对原始token列表去重
如果确认是self.raw_tokens本身有重复条目,可以在循环前直接对列表去重,原有逻辑不需要修改:
# 对原始token去重后再遍历,保留原有顺序 unique_raw_tokens = list(dict.fromkeys(self.raw_tokens)) for rt in unique_raw_tokens: # 原有逻辑保持不变即可 ...
内容的提问来源于stack exchange,提问作者nikom
相关产品推荐
相关产品推荐

