You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C语言实现文件字母统计功能时计数结果偏高问题求助

C语言字母统计代码问题排查

导致统计结果偏高的核心原因

  • 行遍历范围错误:当前内层循环使用pos < sizeof(line)作为终止条件,每次都会遍历完整的200字节缓冲区。但fgets仅会将实际读取的行内容和结尾的\0写入缓冲区,剩余未被写入的空间是未初始化的垃圾值,若垃圾值中存在符合isalpha判断的字符,就会被额外统计,直接导致结果偏高。应将循环条件修改为pos < strlen(line),仅遍历实际读取到的行内容。
  • 计数器数组未初始化:如果调用count_letters函数前,没有将传入的counters数组的26个元素全部初始化为0,数组中的随机初始值会被叠加到统计结果中,也会导致数值偏高。
  • 潜在未定义行为:isalpha、tolower要求输入参数为EOF或者可转换为unsigned char的数值,若系统中char为有符号类型,读取到负值字符时传入这两个函数会触发未定义行为,可能引发异常统计。

修复后的代码

#include <stdio.h>
#include <ctype.h>
#include <string.h>

void count_letters(const char *filename, int counters[26]) {
    // 优先初始化计数器,避免调用方未初始化导致错误
    memset(counters, 0, sizeof(int) * 26);
    FILE* in_file = fopen(filename, "r");
    if(in_file == NULL){
        printf("Error(count_letters): Could not open file %s\n",filename);
        return;
    }
    char line[200];
    while(fgets(line, sizeof(line), in_file) != NULL){
        // 仅遍历当前行实际长度的内容
        for(int pos = 0; pos < strlen(line); pos++){
            unsigned char ch = (unsigned char)line[pos];
            if(isalpha(ch)){
                // 直接计算下标,无需遍历26个字母比对,效率更高
                int idx = tolower(ch) - 'a';
                counters[idx]++;
            }
        }
    }
    fclose(in_file);
    return;
}

内容的提问来源于stack exchange,提问作者user15356836

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 19:06:03