如何将带索引标签的分词元组列表与原字符串对齐生成目标结果
问题:将分词标签元组列表与原字符串对齐并去重
需要将包含字符索引、标签、分词的元组列表idx_tag_token与原字符串word_string对齐,生成由原字符串按空格分割的元素及其对应标签组成的元组列表。现有代码生成的word_tag_list存在重复项,无法得到期望输出,需解决该问题。
数据
word_string = "At London, the 12th in February, 1942, and for that that reason Mark's (3) wins, American parts" idx_tag_token =[(0, 'O', 'At'), (3, 'GPE-B', 'London'), (9, 'O', ','), (11, 'DATE-B', 'the'), (15, 'DATE-I', '12th'), (20, 'O', 'in'), (23, 'DATE-B', 'February'), (31, 'DATE-I', ','), (33, 'DATE-I', '1942'), (37, 'O', ','), (39, 'O', 'and'), (43, 'O', 'for'), (47, 'O', 'that'), (52, 'O', 'that'), (57, 'O', 'reason'), (64, 'PERSON-B', 'Mark'), (68, 'O', "'s"), (71, 'O', '('), (72, 'O', '3'), (73, 'O', ')'), (75, 'O', 'wins'), (79, 'O', ','), (81, 'NORP-B', 'American'), (90, 'O', 'parts')]
原代码
def find_word_from_index(idx, word_string): words = word_string.split() current_index = 0 for word in words: start_index = current_index end_index = current_index + len(word) - 1 if start_index <= idx <= end_index: return word current_index = end_index + 2 return None word_tag_list = [] for index, tag, _ in idx_tag_token: word = find_word_from_index(index, word_string) word_tag_list.append((word, tag)) word_tag_list
当前输出
[('At', 'O'), ('London,', 'GPE-B'), ('London,', 'O'), ('the', 'DATE-B'), ('12th', 'DATE-I'), ('in', 'O'), ('February,', 'DATE-B'), ('February,', 'DATE-I'), ('1942,', 'DATE-I'), ('1942,', 'O'), ('and', 'O'), ('for', 'O'), ('that', 'O'), ('that', 'O'), ('reason', 'O'), ("Mark's", 'PERSON-B'), ("Mark's", 'O'), ('(3)', 'O'), ('(3)', 'O'), ('(3)', 'O'), ('wins,', 'O'), ('wins,', 'O'), ('American', 'NORP-B'), ('parts', 'O')]
期望输出
[('At', 'O'), ('London,', 'GPE-B'), ('the', 'DATE-B'), ('12th', 'DATE-I'), ('in', 'O'), ('February,', 'DATE-B'), ('1942,', 'DATE-I'), ('and', 'O'), ('for', 'O'), ('that', 'O'), ('that', 'O'), ('reason', 'O'), ("Mark's", 'PERSON-B'), ('(3)', 'O'), ('wins,', 'O'), ('American', 'NORP-B'), ('parts', 'O')]
解决方法
问题根源在于idx_tag_token中多个元素对应原字符串中同一个空格分割的单词(比如标点、括号内字符属于同一单词),遍历idx_tag_token会导致同一单词多次出现。正确思路是以原字符串的单词为核心,为每个单词匹配对应标签:
- 先获取原字符串每个空格分割单词的起始、结束字符索引;
- 对每个单词,找到
idx_tag_token中所有索引落在该单词范围内的标签; - 优先选择第一个非'O'的标签,若全为'O'则保留'O'。
修正后代码
def get_word_indices(word_string): words = word_string.split() word_indices = [] current_idx = 0 for word in words: start = current_idx end = current_idx + len(word) - 1 word_indices.append((word, start, end)) current_idx = end + 2 # 加上空格的长度 return word_indices # 生成单词及其索引范围列表 word_indices = get_word_indices(word_string) word_tag_list = [] # 为每个单词匹配标签 for word, start, end in word_indices: # 筛选出当前单词范围内的所有标签 related_tags = [tag for idx, tag, _ in idx_tag_token if start <= idx <= end] # 优先取第一个非'O'的标签,无则取'O' selected_tag = next((t for t in related_tags if t != 'O'), 'O') word_tag_list.append((word, selected_tag)) print(word_tag_list)
输出结果
运行后即可得到与期望一致的无重复元组列表。
内容的提问来源于stack exchange,提问作者doine
相关产品推荐
相关产品推荐

