You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用RE与NLTK实现BIS字符标注:多余空格标注问题求助

修复字符级BIS格式标注中的空格错误

问题描述

我尝试用正则表达式和NLTK实现字符级BIS格式标注:每个字符标注为B(token起始)、I(token中间/末尾)、S(空格)。例如句子"The pen is on the table."的正确标注应该是BIISBIISBISBISBIISBIIIIB,但我的代码输出是BIISBIISBISBISBIISBIIIISB——错误点在于table和句号.之间多了一个S(空格标注),这是因为原代码错误地在这两个无空格分隔的token之间插入了空格对应的S。

原代码如下:

from nltk.tokenize import word_tokenize
import re
p = "The pen is on the table."
# Split text into words using NLTK
text = word_tokenize(p)
print(text)
initial_char = [x.replace(x[0],'B') for x in text]
print(initial_char)
def listToString(s):
    # initialize an empty string
    str1 = " "
    # return string
    return (str1.join(s))
new = listToString(initial_char)
print(new)
def start_from_sec(my_text):
    return ' '.join([f'{word[0]}{(len(word) - 1) * "I"}' for word in my_text.split()])
res = start_from_sec(new)
p = re.sub(' ', 'S', res)
print(p)

错误原因

问题出在两个核心点:

  1. NLTK的word_tokenize会将独立的标点符号(比如这里的.)拆分为单独的token,原文本中table和.之间没有空格,但代码中用空格连接所有token的标注字符串,之后又把所有空格替换成S,导致在本不该有空格的位置插入了S。
  2. 原代码的处理逻辑绕了弯路,通过字符串拼接和替换的方式生成标注,没有考虑token在原文本中的实际分隔情况(是空格分隔还是直接相连)。

修复后的代码

我们可以直接遍历原文本的字符,结合tokenize的结果来精准标注,确保只有原文本中的空格才会被标注为S:

from nltk.tokenize import word_tokenize

p = "The pen is on the table."
tokens = word_tokenize(p)
result = []
current_token_idx = 0
current_char_idx_in_token = 0

for char in p:
    if char == ' ':
        # 空格直接标注为S
        result.append('S')
        # 遇到空格后,下一个字符属于新的token,重置字符索引
        current_char_idx_in_token = 0
    else:
        # 当前字符属于当前token
        if current_char_idx_in_token == 0:
            # token的第一个字符,标注为B
            result.append('B')
        else:
            # token的后续字符,标注为I
            result.append('I')
        current_char_idx_in_token += 1
        # 如果当前token的所有字符都处理完了,切换到下一个token
        if current_char_idx_in_token == len(tokens[current_token_idx]):
            current_token_idx += 1
            current_char_idx_in_token = 0

# 把结果列表拼接成字符串
final_label = ''.join(result)
print(final_label)  # 输出: BIISBIISBISBISBIISBIIIIB

代码解释

  1. 首先用word_tokenize得到token列表,这里是['The', 'pen', 'is', 'on', 'the', 'table', '.']。
  2. 遍历原文本的每个字符:
    • 遇到空格时,直接添加S到结果,同时重置当前token的字符索引,因为下一个字符属于新的token。
    • 非空格字符时,判断是当前token的第一个字符(标注B)还是后续字符(标注I),然后推进字符索引。当当前token的所有字符都处理完,就切换到下一个token。
  3. 最后把结果列表拼接成字符串,得到正确的BIS标注。

这样处理就能精准对应原文本的字符结构,不会在无空格的token之间错误插入S了。

内容的提问来源于stack exchange,提问作者Elias

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 08:37:32