You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

字典单词计数出现Off-by-one错误:实际143091却返回143092

问题:CS50x习题集5单词计数结果偏多

我正在完成CS50x习题集5,编写的代码用于统计字典文件的单词总数:打开文件后用fgets()逐行读取单词,每读取一个有效单词就将count加1。但实际字典共有143091个单词,代码却反复返回143092。用仅含两个单词的小文件测试时,代码能正确计数为2。

代码如下:

#include <ctype.h>
#include <stdbool.h>
#include <stdio.h>
#include <stdlib.h>

int main(void)
{
    // Open files
    FILE *mydictionary = fopen("dictionaries/large", "r");
    if (mydictionary == NULL)
    {
        fclose(mydictionary);
        return 1;
    }

    // Variables
    char word[45];
    int count = 0;

    // Count words
    while (true)
    {
        if (fgets(word, 45, mydictionary) == NULL || isspace(word[0]))
        {
            break;
        }
        count++;
    }

    printf("%i\n", count);
}

问题原因

  1. 长单词被拆分计数:代码用fgets(word, 45, ...)读取内容,最多只能读取44个字符加终止符\0。如果字典里存在长度超过44的单词,fgets会将这个单词拆分成多次读取,每次读取的内容开头都不是空白字符,因此会被多次计数,导致总数多1。小测试文件里的单词长度都小于45,不会触发这个问题。
  2. 未定义行为风险:当fgets返回NULL(文件读取完毕或出错)时,仍访问word[0]调用isspace,属于未定义行为,可能导致不可预期的结果。
  3. 空白行处理错误:代码遇到空白行就直接break终止循环,而非跳过空白行继续统计,若文件中间存在空白行,会导致计数提前终止(不过这不是当前计数偏多的原因,但属于代码bug)。

解决方法

推荐使用fscanf读取单词,它会自动跳过所有空白字符(换行、空格等),并完整读取每个单词,避免长单词拆分的问题,同时简化逻辑:

#include <stdio.h>
#include <stdlib.h>

int main(void)
{
    FILE *mydictionary = fopen("dictionaries/large", "r");
    if (mydictionary == NULL)
    {
        // 文件打开失败时不需要fclose,因为指针是NULL
        return 1;
    }

    char word[45];
    int count = 0;
    // %44s 限制读取长度,避免缓冲区溢出
    while (fscanf(mydictionary, "%44s", word) != EOF)
    {
        count++;
    }

    printf("%i\n", count);
    fclose(mydictionary); // 记得关闭文件
    return 0;
}

如果坚持使用fgets,需要修正逻辑:先判断fgets的返回值,再用continue跳过空白行,同时处理长单词的剩余字符:

#include <ctype.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>

int main(void)
{
    FILE *mydictionary = fopen("dictionaries/large", "r");
    if (mydictionary == NULL)
    {
        return 1;
    }

    char word[45];
    int count = 0;

    while (fgets(word, 45, mydictionary) != NULL)
    {
        // 跳过空白行
        if (isspace((unsigned char)word[0]))
        {
            continue;
        }
        count++;
        // 如果单词过长,跳过剩余字符直到换行或EOF
        if (!strchr(word, '\n'))
        {
            int c;
            while ((c = fgetc(mydictionary)) != EOF && c != '\n');
        }
    }

    printf("%i\n", count);
    fclose(mydictionary);
    return 0;
}

内容的提问来源于stack exchange,提问作者Ratthew

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.02 02:57:27