如何为生成器函数添加基于前序输出的条件?解决括号内标签分配问题
问题:匹配分词标签与原始字符串的括号场景处理
需要将包含字符索引、标签、分词的元组列表id_label_token与原始字符串string进行匹配,现有生成器代码无法处理原始字符串中括号内带有标签的token场景,需实现逻辑对比前后token_type,避免括号内token的标签被遗漏分配。
数据
id_label_token = [(0, 'O', '('), (1, 'DATE-B', '6'), (2, 'DATE-I', ')'), (4, 'DATE-I', '13th'), (9, 'DATE-B', 'February'), (18, 'DATE-I', '1942'), (23, 'O', '('), (24, 'GPE-B', 'N.S.'), (28, 'O', ')')] string = "(6) 13th February 1942 (N.S.)"
现有代码
def get_tokens(tokens): it = iter(tokens) _, token_type, next_token = next(it) word = yield while True: if next_token == word: word = yield next_token, token_type _, token_type, next_token = next(it) else: _, _, tmp = next(it) next_token += tmp it = get_tokens(id_label_token) next(it) out = [it.send(w) for w in string.split()] print(out)
当前输出
[('(6)', 'O'), ('13th', 'DATE-I'), ('February', 'DATE-B'), ('1942', 'DATE-I'), ('(N.S.)', 'O')]
预期输出
[('(6)', 'DATE-B'), ('13th', 'DATE-I'), ('February', 'DATE-B'), ('1942', 'DATE-I'), ('(N.S.)', 'GPE-B')]
解决方案
原代码的核心问题是拼接多个token时,只保留了第一个token的类型,导致括号这类辅助符号的O标签覆盖了内部有效标签。修改思路是:
- 拼接token时同步收集所有对应的标签类型
- 匹配到完整单词后,从收集的标签中选取有效标签:优先取非
O且以-B结尾的起始标签,若没有则取非O的-I标签,最后才用O
修改后的代码:
def get_tokens(tokens): it = iter(tokens) _, current_type, current_token = next(it) collected_types = [current_type] # 收集当前拼接块的所有标签 word = yield while True: if current_token == word: # 从收集的标签中选有效类型:优先-B,再-I,最后O selected_type = 'O' for t in collected_types: if t != 'O': if t.endswith('-B'): selected_type = t break elif t.endswith('-I') and selected_type == 'O': selected_type = t # 返回结果,重置收集器 word = yield current_token, selected_type _, current_type, current_token = next(it) collected_types = [current_type] else: # 继续拼接token,同时收集标签 _, next_type, tmp = next(it) current_token += tmp collected_types.append(next_type) it = get_tokens(id_label_token) next(it) out = [it.send(w) for w in string.split()] print(out)
运行后输出与预期一致:
[('(6)', 'DATE-B'), ('13th', 'DATE-I'), ('February', 'DATE-B'), ('1942', 'DATE-I'), ('(N.S.)', 'GPE-B')]
内容的提问来源于stack exchange,提问作者doine
相关产品推荐
相关产品推荐

