使用RE与NLTK实现BIS字符标注:多余空格标注问题求助
修复字符级BIS格式标注中的空格错误
问题描述
我尝试用正则表达式和NLTK实现字符级BIS格式标注:每个字符标注为B(token起始)、I(token中间/末尾)、S(空格)。例如句子"The pen is on the table."的正确标注应该是BIISBIISBISBISBIISBIIIIB,但我的代码输出是BIISBIISBISBISBIISBIIIISB——错误点在于table和句号.之间多了一个S(空格标注),这是因为原代码错误地在这两个无空格分隔的token之间插入了空格对应的S。
原代码如下:
from nltk.tokenize import word_tokenize import re p = "The pen is on the table." # Split text into words using NLTK text = word_tokenize(p) print(text) initial_char = [x.replace(x[0],'B') for x in text] print(initial_char) def listToString(s): # initialize an empty string str1 = " " # return string return (str1.join(s)) new = listToString(initial_char) print(new) def start_from_sec(my_text): return ' '.join([f'{word[0]}{(len(word) - 1) * "I"}' for word in my_text.split()]) res = start_from_sec(new) p = re.sub(' ', 'S', res) print(p)
错误原因
问题出在两个核心点:
- NLTK的
word_tokenize会将独立的标点符号(比如这里的.)拆分为单独的token,原文本中table和.之间没有空格,但代码中用空格连接所有token的标注字符串,之后又把所有空格替换成S,导致在本不该有空格的位置插入了S。 - 原代码的处理逻辑绕了弯路,通过字符串拼接和替换的方式生成标注,没有考虑token在原文本中的实际分隔情况(是空格分隔还是直接相连)。
修复后的代码
我们可以直接遍历原文本的字符,结合tokenize的结果来精准标注,确保只有原文本中的空格才会被标注为S:
from nltk.tokenize import word_tokenize p = "The pen is on the table." tokens = word_tokenize(p) result = [] current_token_idx = 0 current_char_idx_in_token = 0 for char in p: if char == ' ': # 空格直接标注为S result.append('S') # 遇到空格后,下一个字符属于新的token,重置字符索引 current_char_idx_in_token = 0 else: # 当前字符属于当前token if current_char_idx_in_token == 0: # token的第一个字符,标注为B result.append('B') else: # token的后续字符,标注为I result.append('I') current_char_idx_in_token += 1 # 如果当前token的所有字符都处理完了,切换到下一个token if current_char_idx_in_token == len(tokens[current_token_idx]): current_token_idx += 1 current_char_idx_in_token = 0 # 把结果列表拼接成字符串 final_label = ''.join(result) print(final_label) # 输出: BIISBIISBISBISBIISBIIIIB
代码解释
- 首先用
word_tokenize得到token列表,这里是['The', 'pen', 'is', 'on', 'the', 'table', '.']。 - 遍历原文本的每个字符:
- 遇到空格时,直接添加S到结果,同时重置当前token的字符索引,因为下一个字符属于新的token。
- 非空格字符时,判断是当前token的第一个字符(标注B)还是后续字符(标注I),然后推进字符索引。当当前token的所有字符都处理完,就切换到下一个token。
- 最后把结果列表拼接成字符串,得到正确的BIS标注。
这样处理就能精准对应原文本的字符结构,不会在无空格的token之间错误插入S了。
内容的提问来源于stack exchange,提问作者Elias
相关产品推荐
相关产品推荐

