You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python tokenize分词函数处理全空格输入返回空串问题排查

Python词频统计tokenize函数实现与修复

需求背景

需要使用Python开发词频统计程序,统计指定文本的单词类型与出现频率,要求不计入停用词、空格以及+-??:";等特殊字符。程序第一步需实现tokenize分词函数,对应测试用例如下:

if hasattr(wordfreq, "tokenize"):
    fun_count = fun_count + 1
    test(wordfreq.tokenize, [], [])
    test(wordfreq.tokenize, [""], [])
    test(wordfreq.tokenize, ["   "], [])
    test(wordfreq.tokenize, ["This is a simple sentence"], ["this","is","a","simple","sentence"])
    test(wordfreq.tokenize, ["I told you!"], ["i","told","you","!"])
    test(wordfreq.tokenize, ["The 10 little chicks"], ["the","10","little","chicks"])
    test(wordfreq.tokenize, ["15th anniversary"], ["15","th","anniversary"])
    test(wordfreq.tokenize, ["He is in the room, she said."], ["he","is","in","the","room",",","she","said","."])
else:
    print("tokenize is not implemented yet!")

问题现象

已编写的tokenize函数8个测试仅通过7个,报错显示tokenize([' ']) == []条件不成立,实际返回结果为['']。
原有代码如下:

def tokenize(lines):
    words = []
    for line in lines:
        start = 0
        while start < len(line):
            while start < len(line) and line[start].isspace():
                start = start + 1
            end = start
            if end < len(line) and line[end].isdigit():
                end = start
                while end < len(line) and line[end].isdigit():
                    end = end + 1
                words.append(line[start:end])
                start = end
            elif end < len(line) and line[end].isalpha():
                end = start
                while end < len(line) and line[end].isalpha():
                    end = end + 1
                words.append(line[start:end].lower())
                start = end
            else: 
                end = start
                end < len(line)
                end = end + 1
                words.append(line[start:end])
                start = end 
    return words

问题原因

  1. 跳过空格的循环执行完成后,没有判断start是否已经走到行尾,直接进入了后续分词逻辑
  2. 当输入为全空格字符串时,跳过空格的循环结束后start等于字符串长度,既不满足数字分支也不满足字母分支,直接进入else分支
  3. else分支未做边界判断,直接执行end+1后切片,当start等于字符串长度时,line[start:end]会得到空字符串,被添加到结果列表中导致报错
  4. else分支中end < len(line)是无效表达式,没有任何判断作用,属于冗余错误代码

修复方案

在跳过空格的逻辑后新增边界判断,如果start已经超出字符串索引范围,直接结束当前行的处理,不再执行后续分词逻辑,同时清理原有代码中的冗余赋值和无效表达式。

修正后代码

def tokenize(lines):
    words = []
    for line in lines:
        start = 0
        while start < len(line):
            # 跳过所有空格
            while start < len(line) and line[start].isspace():
                start += 1
            # 新增边界判断:已经走到行尾直接结束
            if start >= len(line):
                break
            end = start
            if line[end].isdigit():
                while end < len(line) and line[end].isdigit():
                    end += 1
                words.append(line[start:end])
                start = end
            elif line[end].isalpha():
                while end < len(line) and line[end].isalpha():
                    end += 1
                words.append(line[start:end].lower())
                start = end
            else:
                # 单个非字母数字非空格字符直接返回
                end += 1
                words.append(line[start:end])
                start = end
    return words

代码差异说明

  • 新增空格跳过逻辑后的边界判断,避免全空格输入时产生空字符串结果
  • 清理了数字、字母分支中冗余的end = start重复赋值
  • 移除了else分支中无效的end < len(line)表达式和冗余的end = start赋值
  • 简化了数字、字母分支的判断逻辑,因为已提前做了边界校验,无需重复判断end < len(line)

内容的提问来源于stack exchange,提问作者Aut_K_H

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 01:09:00