字典单词计数出现Off-by-one错误:实际143091却返回143092
问题:CS50x习题集5单词计数结果偏多
我正在完成CS50x习题集5,编写的代码用于统计字典文件的单词总数:打开文件后用fgets()逐行读取单词,每读取一个有效单词就将count加1。但实际字典共有143091个单词,代码却反复返回143092。用仅含两个单词的小文件测试时,代码能正确计数为2。
代码如下:
#include <ctype.h> #include <stdbool.h> #include <stdio.h> #include <stdlib.h> int main(void) { // Open files FILE *mydictionary = fopen("dictionaries/large", "r"); if (mydictionary == NULL) { fclose(mydictionary); return 1; } // Variables char word[45]; int count = 0; // Count words while (true) { if (fgets(word, 45, mydictionary) == NULL || isspace(word[0])) { break; } count++; } printf("%i\n", count); }
问题原因
- 长单词被拆分计数:代码用
fgets(word, 45, ...)读取内容,最多只能读取44个字符加终止符\0。如果字典里存在长度超过44的单词,fgets会将这个单词拆分成多次读取,每次读取的内容开头都不是空白字符,因此会被多次计数,导致总数多1。小测试文件里的单词长度都小于45,不会触发这个问题。 - 未定义行为风险:当
fgets返回NULL(文件读取完毕或出错)时,仍访问word[0]调用isspace,属于未定义行为,可能导致不可预期的结果。 - 空白行处理错误:代码遇到空白行就直接
break终止循环,而非跳过空白行继续统计,若文件中间存在空白行,会导致计数提前终止(不过这不是当前计数偏多的原因,但属于代码bug)。
解决方法
推荐使用fscanf读取单词,它会自动跳过所有空白字符(换行、空格等),并完整读取每个单词,避免长单词拆分的问题,同时简化逻辑:
#include <stdio.h> #include <stdlib.h> int main(void) { FILE *mydictionary = fopen("dictionaries/large", "r"); if (mydictionary == NULL) { // 文件打开失败时不需要fclose,因为指针是NULL return 1; } char word[45]; int count = 0; // %44s 限制读取长度,避免缓冲区溢出 while (fscanf(mydictionary, "%44s", word) != EOF) { count++; } printf("%i\n", count); fclose(mydictionary); // 记得关闭文件 return 0; }
如果坚持使用fgets,需要修正逻辑:先判断fgets的返回值,再用continue跳过空白行,同时处理长单词的剩余字符:
#include <ctype.h> #include <stdio.h> #include <stdlib.h> #include <string.h> int main(void) { FILE *mydictionary = fopen("dictionaries/large", "r"); if (mydictionary == NULL) { return 1; } char word[45]; int count = 0; while (fgets(word, 45, mydictionary) != NULL) { // 跳过空白行 if (isspace((unsigned char)word[0])) { continue; } count++; // 如果单词过长,跳过剩余字符直到换行或EOF if (!strchr(word, '\n')) { int c; while ((c = fgetc(mydictionary)) != EOF && c != '\n'); } } printf("%i\n", count); fclose(mydictionary); return 0; }
内容的提问来源于stack exchange,提问作者Ratthew
相关产品推荐
相关产品推荐

