You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为生成器函数添加基于前序输出的条件?解决括号内标签分配问题

问题:匹配分词标签与原始字符串的括号场景处理

需要将包含字符索引、标签、分词的元组列表id_label_token与原始字符串string进行匹配,现有生成器代码无法处理原始字符串中括号内带有标签的token场景,需实现逻辑对比前后token_type,避免括号内token的标签被遗漏分配。

数据

id_label_token = [(0, 'O', '('),
                  (1, 'DATE-B', '6'),
                  (2, 'DATE-I', ')'),
                  (4, 'DATE-I', '13th'),
                  (9, 'DATE-B', 'February'),
                  (18, 'DATE-I', '1942'),
                  (23, 'O', '('),
                  (24, 'GPE-B', 'N.S.'),
                  (28, 'O', ')')]

string = "(6) 13th February 1942 (N.S.)"

现有代码

def get_tokens(tokens):
    it = iter(tokens)
    _, token_type, next_token = next(it)
    word = yield
    while True:
        if next_token == word:
            word = yield next_token, token_type
            _, token_type, next_token = next(it)
        else:
            _, _, tmp = next(it)
            next_token += tmp

it = get_tokens(id_label_token)
next(it)
out = [it.send(w) for w in string.split()]
print(out)

当前输出

[('(6)', 'O'), ('13th', 'DATE-I'), ('February', 'DATE-B'), ('1942', 'DATE-I'), ('(N.S.)', 'O')]

预期输出

[('(6)', 'DATE-B'), ('13th', 'DATE-I'), ('February', 'DATE-B'), ('1942', 'DATE-I'), ('(N.S.)', 'GPE-B')]

解决方案

原代码的核心问题是拼接多个token时,只保留了第一个token的类型,导致括号这类辅助符号的O标签覆盖了内部有效标签。修改思路是:

  • 拼接token时同步收集所有对应的标签类型
  • 匹配到完整单词后,从收集的标签中选取有效标签:优先取非O且以-B结尾的起始标签,若没有则取非O的-I标签,最后才用O

修改后的代码:

def get_tokens(tokens):
    it = iter(tokens)
    _, current_type, current_token = next(it)
    collected_types = [current_type]  # 收集当前拼接块的所有标签
    word = yield
    while True:
        if current_token == word:
            # 从收集的标签中选有效类型:优先-B,再-I,最后O
            selected_type = 'O'
            for t in collected_types:
                if t != 'O':
                    if t.endswith('-B'):
                        selected_type = t
                        break
                    elif t.endswith('-I') and selected_type == 'O':
                        selected_type = t
            # 返回结果,重置收集器
            word = yield current_token, selected_type
            _, current_type, current_token = next(it)
            collected_types = [current_type]
        else:
            # 继续拼接token,同时收集标签
            _, next_type, tmp = next(it)
            current_token += tmp
            collected_types.append(next_type)

it = get_tokens(id_label_token)
next(it)
out = [it.send(w) for w in string.split()]
print(out)

运行后输出与预期一致:

[('(6)', 'DATE-B'), ('13th', 'DATE-I'), ('February', 'DATE-B'), ('1942', 'DATE-I'), ('(N.S.)', 'GPE-B')]

内容的提问来源于stack exchange,提问作者doine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 10:17:37