You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python列表元素重复记录 如何实现概念同义词打印结果去重

同义词提取功能输出重复问题解决

问题描述

开发从数据集/样本中提取同义词功能时,使用现有代码打印概念与关联同义词会重复输出两次,可读性差,需要实现仅打印唯一结果。

现有问题代码

for rt in self.raw_tokens:
    concept = None
    if rt.startswith('<'):
        # 如果是概念(实体)
        if 'value' in ElementTree.fromstring(rt).attrib:
            string_token = ElementTree.fromstring(rt).text
            concept = ElementTree.fromstring(rt).tag
        # 如果是同义词,直接取文本并移除样本中的标签
        else:
            string_token = ElementTree.fromstring(rt).text
            #print('Token: [' + rt + ']')
            print("Concepts & Synonyms are present in the sample :<%s> = %s" %(ElementTree.fromstring(rt).tag,string_token))
            #print('CONCEPT: ' + concept)
    else:
            string_token = rt
    toks.append((concept, string_token))

输出对比

预期输出

Concepts & Synonyms are present in the sample :<<I_LOVE>I like</I_LOVE>> = I like
Concepts & Synonyms are present in the sample :<<I_WANTS>NEED</I_WANTS>> = NEED
Concepts & Synonyms are present in the sample :<<NEW_YORK>NEW YORK</NEW_YORK>> = NEW YORK
Concepts & Synonyms are present in the sample :<<I_WANTS>wish</I_WANTS>> = wish
Concepts & Synonyms are present in the sample :<<NEW_YORK>BIG APPLE</NEW_YORK>> = BIG APPLE

实际输出(存在重复)

Concepts & Synonyms are present in the sample :<<I_LOVE>I like</I_LOVE>> = I like
Concepts & Synonyms are present in the sample :<<I_WANTS>NEED</I_WANTS>> = NEED
Concepts & Synonyms are present in the sample :<<NEW_YORK>NEW YORK</NEW_YORK>> = NEW YORK
Concepts & Synonyms are present in the sample :<<I_WANTS>wish</I_WANTS>> = wish
Concepts & Synonyms are present in the sample :<<NEW_YORK>BIG APPLE</NEW_YORK>> = BIG APPLE
Concepts & Synonyms are present in the sample :<<I_LOVE>I like</I_LOVE>> = I like
Concepts & Synonyms are present in the sample :<<I_WANTS>NEED</I_WANTS>> = NEED
Concepts & Synonyms are present in the sample :<<NEW_YORK>NEW YORK</NEW_YORK>> = NEW YORK
Concepts & Synonyms are present in the sample :<<I_WANTS>wish</I_WANTS>> = wish
Concepts & Synonyms are present in the sample :<<NEW_YORK>BIG APPLE</NEW_YORK>> = BIG APPLE

解决方案

输出重复的核心原因是self.raw_tokens本身存在重复条目,或是该段循环代码被重复调用,两种修复方案如下:

方案1:新增已打印集合去重(对原有代码改动最小)

在循环外初始化空集合记录已打印的内容,打印前先判断是否已存在,不存在才打印同时加入集合,同时优化重复解析XML的冗余逻辑:

# 初始化去重集合,放在循环外部
printed_entries = set()

for rt in self.raw_tokens:
    concept = None
    if rt.startswith('<'):
        # 提前解析XML避免重复调用
        elem = ElementTree.fromstring(rt)
        if 'value' in elem.attrib:
            string_token = elem.text
            concept = elem.tag
        else:
            string_token = elem.text
            tag = elem.tag
            # 用概念+同义词的组合作为唯一判断键
            entry_key = (tag, string_token)
            if entry_key not in printed_entries:
                printed_entries.add(entry_key)
                print(f"Concepts & Synonyms are present in the sample :<{tag}> = {string_token}")
    else:
            string_token = rt
    toks.append((concept, string_token))

方案2:提前对原始token列表去重

如果确认是self.raw_tokens本身有重复条目,可以在循环前直接对列表去重,原有逻辑不需要修改:

# 对原始token去重后再遍历,保留原有顺序
unique_raw_tokens = list(dict.fromkeys(self.raw_tokens))
for rt in unique_raw_tokens:
    # 原有逻辑保持不变即可
    ...

内容的提问来源于stack exchange,提问作者nikom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 04:06:01