You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将带索引标签的分词元组列表与原字符串对齐生成目标结果

问题:将分词标签元组列表与原字符串对齐并去重

需要将包含字符索引、标签、分词的元组列表idx_tag_token与原字符串word_string对齐,生成由原字符串按空格分割的元素及其对应标签组成的元组列表。现有代码生成的word_tag_list存在重复项,无法得到期望输出,需解决该问题。

数据

word_string = "At London, the 12th in February, 1942, and for that that reason Mark's (3) wins, American parts"

idx_tag_token =[(0, 'O', 'At'),
                (3, 'GPE-B', 'London'),
                (9, 'O', ','),
                (11, 'DATE-B', 'the'),
                (15, 'DATE-I', '12th'),
                (20, 'O', 'in'),
                (23, 'DATE-B', 'February'),
                (31, 'DATE-I', ','),
                (33, 'DATE-I', '1942'),
                (37, 'O', ','),
                (39, 'O', 'and'),
                (43, 'O', 'for'),
                (47, 'O', 'that'),
                (52, 'O', 'that'),
                (57, 'O', 'reason'),
                (64, 'PERSON-B', 'Mark'),
                (68, 'O', "'s"),
                (71, 'O', '('),
                (72, 'O', '3'),
                (73, 'O', ')'),
                (75, 'O', 'wins'),
                (79, 'O', ','),
                (81, 'NORP-B', 'American'),
                (90, 'O', 'parts')]

原代码

def find_word_from_index(idx, word_string):
    words = word_string.split()
    current_index = 0

    for word in words:
        start_index = current_index
        end_index = current_index + len(word) - 1
        if start_index <= idx <= end_index:
            return word
        current_index = end_index + 2
    return None


word_tag_list = []
for index, tag, _ in idx_tag_token:
    word = find_word_from_index(index, word_string)
    word_tag_list.append((word, tag))
word_tag_list

当前输出

[('At', 'O'),
 ('London,', 'GPE-B'),
 ('London,', 'O'),
 ('the', 'DATE-B'),
 ('12th', 'DATE-I'),
 ('in', 'O'),
 ('February,', 'DATE-B'),
 ('February,', 'DATE-I'),
 ('1942,', 'DATE-I'),
 ('1942,', 'O'),
 ('and', 'O'),
 ('for', 'O'),
 ('that', 'O'),
 ('that', 'O'),
 ('reason', 'O'),
 ("Mark's", 'PERSON-B'),
 ("Mark's", 'O'),
 ('(3)', 'O'),
 ('(3)', 'O'),
 ('(3)', 'O'),
 ('wins,', 'O'),
 ('wins,', 'O'),
 ('American', 'NORP-B'),
 ('parts', 'O')]

期望输出

[('At', 'O'),
('London,', 'GPE-B'),
('the', 'DATE-B'),
('12th', 'DATE-I'),
('in', 'O'),
('February,', 'DATE-B'),
('1942,', 'DATE-I'),
('and', 'O'),
('for', 'O'),
('that', 'O'),
('that', 'O'),
('reason', 'O'),
("Mark's", 'PERSON-B'),
('(3)', 'O'),
('wins,', 'O'),
('American', 'NORP-B'),
('parts', 'O')]

解决方法

问题根源在于idx_tag_token中多个元素对应原字符串中同一个空格分割的单词(比如标点、括号内字符属于同一单词),遍历idx_tag_token会导致同一单词多次出现。正确思路是以原字符串的单词为核心,为每个单词匹配对应标签:

  1. 先获取原字符串每个空格分割单词的起始、结束字符索引;
  2. 对每个单词,找到idx_tag_token中所有索引落在该单词范围内的标签;
  3. 优先选择第一个非'O'的标签,若全为'O'则保留'O'。

修正后代码

def get_word_indices(word_string):
    words = word_string.split()
    word_indices = []
    current_idx = 0
    for word in words:
        start = current_idx
        end = current_idx + len(word) - 1
        word_indices.append((word, start, end))
        current_idx = end + 2  # 加上空格的长度
    return word_indices

# 生成单词及其索引范围列表
word_indices = get_word_indices(word_string)
word_tag_list = []

# 为每个单词匹配标签
for word, start, end in word_indices:
    # 筛选出当前单词范围内的所有标签
    related_tags = [tag for idx, tag, _ in idx_tag_token if start <= idx <= end]
    # 优先取第一个非'O'的标签,无则取'O'
    selected_tag = next((t for t in related_tags if t != 'O'), 'O')
    word_tag_list.append((word, selected_tag))

print(word_tag_list)

输出结果

运行后即可得到与期望一致的无重复元组列表。

内容的提问来源于stack exchange,提问作者doine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 06:12:04