Python tokenize分词函数处理全空格输入返回空串问题排查
Python词频统计tokenize函数实现与修复
需求背景
需要使用Python开发词频统计程序,统计指定文本的单词类型与出现频率,要求不计入停用词、空格以及+-??:";等特殊字符。程序第一步需实现tokenize分词函数,对应测试用例如下:
if hasattr(wordfreq, "tokenize"): fun_count = fun_count + 1 test(wordfreq.tokenize, [], []) test(wordfreq.tokenize, [""], []) test(wordfreq.tokenize, [" "], []) test(wordfreq.tokenize, ["This is a simple sentence"], ["this","is","a","simple","sentence"]) test(wordfreq.tokenize, ["I told you!"], ["i","told","you","!"]) test(wordfreq.tokenize, ["The 10 little chicks"], ["the","10","little","chicks"]) test(wordfreq.tokenize, ["15th anniversary"], ["15","th","anniversary"]) test(wordfreq.tokenize, ["He is in the room, she said."], ["he","is","in","the","room",",","she","said","."]) else: print("tokenize is not implemented yet!")
问题现象
已编写的tokenize函数8个测试仅通过7个,报错显示tokenize([' ']) == []条件不成立,实际返回结果为['']。
原有代码如下:
def tokenize(lines): words = [] for line in lines: start = 0 while start < len(line): while start < len(line) and line[start].isspace(): start = start + 1 end = start if end < len(line) and line[end].isdigit(): end = start while end < len(line) and line[end].isdigit(): end = end + 1 words.append(line[start:end]) start = end elif end < len(line) and line[end].isalpha(): end = start while end < len(line) and line[end].isalpha(): end = end + 1 words.append(line[start:end].lower()) start = end else: end = start end < len(line) end = end + 1 words.append(line[start:end]) start = end return words
问题原因
- 跳过空格的循环执行完成后,没有判断
start是否已经走到行尾,直接进入了后续分词逻辑 - 当输入为全空格字符串时,跳过空格的循环结束后
start等于字符串长度,既不满足数字分支也不满足字母分支,直接进入else分支 - else分支未做边界判断,直接执行
end+1后切片,当start等于字符串长度时,line[start:end]会得到空字符串,被添加到结果列表中导致报错 - else分支中
end < len(line)是无效表达式,没有任何判断作用,属于冗余错误代码
修复方案
在跳过空格的逻辑后新增边界判断,如果start已经超出字符串索引范围,直接结束当前行的处理,不再执行后续分词逻辑,同时清理原有代码中的冗余赋值和无效表达式。
修正后代码
def tokenize(lines): words = [] for line in lines: start = 0 while start < len(line): # 跳过所有空格 while start < len(line) and line[start].isspace(): start += 1 # 新增边界判断:已经走到行尾直接结束 if start >= len(line): break end = start if line[end].isdigit(): while end < len(line) and line[end].isdigit(): end += 1 words.append(line[start:end]) start = end elif line[end].isalpha(): while end < len(line) and line[end].isalpha(): end += 1 words.append(line[start:end].lower()) start = end else: # 单个非字母数字非空格字符直接返回 end += 1 words.append(line[start:end]) start = end return words
代码差异说明
- 新增空格跳过逻辑后的边界判断,避免全空格输入时产生空字符串结果
- 清理了数字、字母分支中冗余的
end = start重复赋值 - 移除了else分支中无效的
end < len(line)表达式和冗余的end = start赋值 - 简化了数字、字母分支的判断逻辑,因为已提前做了边界校验,无需重复判断
end < len(line)
内容的提问来源于stack exchange,提问作者Aut_K_H
相关产品推荐
相关产品推荐

