You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按指定标点分词?NLTK sent_tokenize使用遇阻求助

解决NLTK分词不符合预期的问题

你的代码存在的核心问题

  • 函数参数被硬覆盖:tokenise函数内部直接重新赋值了input和expected_output,完全忽略了传入的测试用例参数,所有调用都只处理同一个固定字符串。
  • sent_tokenize逻辑不匹配需求:sent_tokenize默认按**句子边界(!、?、.)**分句,不会将逗号作为分隔符,而你的需求是要把,!?.这四类标点都作为拆分依据。
  • 变量未定义+作用域错误:tcase1、tans1等变量未定义就直接调用,且函数内部的result、expected_output是局部变量,外部print语句无法访问,会直接报错。
  • 缩进错误:函数内的print语句未缩进,不会在函数调用时执行。

修正后的代码实现

from nltk.tokenize import RegexpTokenizer

def tokenise(input_str, expected_output):
    # 自定义正则规则:匹配所有非!、?、.的字符序列,处理连续标点场景
    tokenizer = RegexpTokenizer(r'[^!?.]+')
    # 清洗结果:去除空字符串和首尾空格
    result = [s.strip() for s in tokenizer.tokenize(input_str) if s.strip()]
    print('Pass' if result == expected_output else 'Failed!')

# 定义所有测试用例
tcase1 = "Excuse me, where can I find a chicken rice shop?"
tans1 = ["Excuse me", "where can I find a chicken rice shop"]

tcase2 = "OMG!!! It is Friday....where should we go for dinner?"
tans2 = ["OMG", "It is Friday", "where should we go for dinner"]

tcase3 = "He’s nervous, but on the surface he looks calm and ready."
tans3 = ["He’s nervous", "but on the surface he looks calm and ready"]

# 执行测试
tokenise(tcase1, tans1)
tokenise(tcase2, tans2)
tokenise(tcase3, tans3)

关键修改说明

  • 自定义分词规则:用RegexpTokenizer实现按,!?.拆分的需求,正则r'[^!?.]+'会匹配所有不包含目标标点的连续字符,自动跳过任意数量的连续标点。
  • 结果清洗:通过列表推导式过滤拆分后产生的空内容,同时去除每个片段的首尾空格,确保结果和预期完全对齐。
  • 参数正常传递:不再覆盖传入的测试参数,直接使用外部定义的用例数据。
  • 补全变量定义:提前定义好所有测试用例的输入和预期输出,避免未定义错误。

内容的提问来源于stack exchange,提问作者jabs93

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 16:31:02