C语言实现文件字母统计功能时计数结果偏高问题求助
C语言字母统计代码问题排查
导致统计结果偏高的核心原因
- 行遍历范围错误:当前内层循环使用
pos < sizeof(line)作为终止条件,每次都会遍历完整的200字节缓冲区。但fgets仅会将实际读取的行内容和结尾的\0写入缓冲区,剩余未被写入的空间是未初始化的垃圾值,若垃圾值中存在符合isalpha判断的字符,就会被额外统计,直接导致结果偏高。应将循环条件修改为pos < strlen(line),仅遍历实际读取到的行内容。 - 计数器数组未初始化:如果调用
count_letters函数前,没有将传入的counters数组的26个元素全部初始化为0,数组中的随机初始值会被叠加到统计结果中,也会导致数值偏高。 - 潜在未定义行为:
isalpha、tolower要求输入参数为EOF或者可转换为unsigned char的数值,若系统中char为有符号类型,读取到负值字符时传入这两个函数会触发未定义行为,可能引发异常统计。
修复后的代码
#include <stdio.h> #include <ctype.h> #include <string.h> void count_letters(const char *filename, int counters[26]) { // 优先初始化计数器,避免调用方未初始化导致错误 memset(counters, 0, sizeof(int) * 26); FILE* in_file = fopen(filename, "r"); if(in_file == NULL){ printf("Error(count_letters): Could not open file %s\n",filename); return; } char line[200]; while(fgets(line, sizeof(line), in_file) != NULL){ // 仅遍历当前行实际长度的内容 for(int pos = 0; pos < strlen(line); pos++){ unsigned char ch = (unsigned char)line[pos]; if(isalpha(ch)){ // 直接计算下标,无需遍历26个字母比对,效率更高 int idx = tolower(ch) - 'a'; counters[idx]++; } } } fclose(in_file); return; }
内容的提问来源于stack exchange,提问作者user15356836
相关产品推荐
相关产品推荐

