You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

两段Python文本单词统计代码结果差异的原因解析

单词统计代码结果差异解析

两段代码的核心差异在于统计逻辑的粒度不同:

第一段代码逻辑

这段代码先将文本转小写,再用split()分割成独立单词列表,之后统计目标单词在列表中的出现次数:

def popular_words(text:str, list:list)-> dict:
    text=text.lower()
    splited_text=text.split()
    answer={}

    for word in list:
        answer[word]=splited_text.count(word) 

    return answer

print(popular_words('''
When I was One 
I had just begun 
When I was Two 
I was nearly new''', ['i', 'was', 'three', 'near']))

运行结果:{'i': 4, 'was': 3, 'three': 0, 'near': 0}

split()默认按空白字符分割,分割后的单词列表为:
['when', 'i', 'was', 'one', 'i', 'had', 'just', 'begun', 'when', 'i', 'was', 'two', 'i', 'was', 'nearly', 'new']
列表中没有独立的near单词,因此计数为0。

第二段代码逻辑

这段代码转小写后直接用字符串的count()方法统计目标子串的出现次数:

def popular_words(text:str, list:list)-> dict:
    text=text.lower()
    answer={}
    for word in list:
    
        answer[word]=text.count(word)

    return answer


print(popular_words('''
When I was One 
I had just begun 
When I was Two 
I was nearly new''', ['i', 'was', 'three', 'near']))

运行结果:{'i': 4, 'was': 3, 'three': 0, 'near': 1}

字符串的count()方法会匹配所有包含目标子串的位置,不区分是否为独立单词。文本转小写后包含nearly这个单词,其中的near子串会被统计到,因此text.count('near')结果为1。

总结

  • 第一段代码统计的是独立完整单词的出现次数,只有当目标是完全匹配的独立单词时才会被计数
  • 第二段代码统计的是任意位置子串的出现次数,只要文本中存在连续的目标字符序列就会被计数

内容的提问来源于stack exchange,提问作者bangtae

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 02:48:24