如何按指定标点分词?NLTK sent_tokenize使用遇阻求助
解决NLTK分词不符合预期的问题
你的代码存在的核心问题
- 函数参数被硬覆盖:
tokenise函数内部直接重新赋值了input和expected_output,完全忽略了传入的测试用例参数,所有调用都只处理同一个固定字符串。 - sent_tokenize逻辑不匹配需求:
sent_tokenize默认按**句子边界(!、?、.)**分句,不会将逗号作为分隔符,而你的需求是要把,!?.这四类标点都作为拆分依据。 - 变量未定义+作用域错误:
tcase1、tans1等变量未定义就直接调用,且函数内部的result、expected_output是局部变量,外部print语句无法访问,会直接报错。 - 缩进错误:函数内的
print语句未缩进,不会在函数调用时执行。
修正后的代码实现
from nltk.tokenize import RegexpTokenizer def tokenise(input_str, expected_output): # 自定义正则规则:匹配所有非!、?、.的字符序列,处理连续标点场景 tokenizer = RegexpTokenizer(r'[^!?.]+') # 清洗结果:去除空字符串和首尾空格 result = [s.strip() for s in tokenizer.tokenize(input_str) if s.strip()] print('Pass' if result == expected_output else 'Failed!') # 定义所有测试用例 tcase1 = "Excuse me, where can I find a chicken rice shop?" tans1 = ["Excuse me", "where can I find a chicken rice shop"] tcase2 = "OMG!!! It is Friday....where should we go for dinner?" tans2 = ["OMG", "It is Friday", "where should we go for dinner"] tcase3 = "He’s nervous, but on the surface he looks calm and ready." tans3 = ["He’s nervous", "but on the surface he looks calm and ready"] # 执行测试 tokenise(tcase1, tans1) tokenise(tcase2, tans2) tokenise(tcase3, tans3)
关键修改说明
- 自定义分词规则:用
RegexpTokenizer实现按,!?.拆分的需求,正则r'[^!?.]+'会匹配所有不包含目标标点的连续字符,自动跳过任意数量的连续标点。 - 结果清洗:通过列表推导式过滤拆分后产生的空内容,同时去除每个片段的首尾空格,确保结果和预期完全对齐。
- 参数正常传递:不再覆盖传入的测试参数,直接使用外部定义的用例数据。
- 补全变量定义:提前定义好所有测试用例的输入和预期输出,避免未定义错误。
内容的提问来源于stack exchange,提问作者jabs93
相关产品推荐
相关产品推荐

