You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复count_word函数统计speech.txt单词次数返回0的问题

排查count_word函数统计单词次数返回0的常见原因

以下是几种导致统计结果为0但实际存在目标单词的典型问题及修复方案:

1. 未处理文本中的标点符号

如果代码直接按空格分割文本,带有逗号/句号的单词(比如"hello,"、"world.")会被视为和目标单词(比如"hello")不同的字符串,导致统计遗漏。

错误示例代码:

def count_word(target_word):
    with open('speech.txt', 'r') as f:
        words = f.read().split()
        return words.count(target_word)

修复方案:先移除文本中的逗号和句号,再进行分割统计:

def count_word(target_word):
    with open('speech.txt', 'r', encoding='utf-8') as f:
        content = f.read()
        # 清除标点符号
        cleaned_content = content.replace(',', '').replace('.', '')
        words = cleaned_content.split()
        return words.count(target_word)

2. 大小写不匹配

如果目标单词和文本中的单词大小写不一致(比如目标是"Hello",文本里是"hello"、"HELLO"),直接精确匹配会遗漏这些情况。

修复方案:统一转换为小写(或大写)后再比较:

def count_word(target_word):
    target_lower = target_word.lower()
    with open('speech.txt', 'r', encoding='utf-8') as f:
        content = f.read().lower()
        cleaned_content = content.replace(',', '').replace('.', '')
        words = cleaned_content.split()
        return words.count(target_lower)

3. 文件读取异常或路径错误

如果代码未处理文件读取异常,比如文件路径错误、编码不匹配导致读取内容为空,也会返回0。

排查点:

  • 确认speech.txt的路径是否正确(比如是否和脚本在同一目录,或使用绝对路径)
  • 打开文件时指定正确编码(比如encoding='utf-8')
  • 临时打印读取到的内容,确认是否成功读取到目标文本

4. 单词匹配未考虑独立边界(可选优化)

如果文本中存在包含目标单词的更长单词(比如目标是"cat",文本里有"category"),直接分割统计可能误判;用正则匹配独立单词更严谨:

正则方案:

import re

def count_word(target_word):
    with open('speech.txt', 'r', encoding='utf-8') as f:
        content = f.read()
        # 匹配独立单词,忽略大小写
        pattern = re.compile(r'\b' + re.escape(target_word) + r'\b', re.IGNORECASE)
        return len(pattern.findall(content))

内容的提问来源于stack exchange,提问作者panda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 11:00:13