You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C语言统计文本文件所有单词出现次数的代码问题求助

问题分析与修复

你的代码存在几个关键问题导致无法正确统计单词出现次数:

  • 文件读取逻辑错误:第一次用fscanf读取一个单词后,紧接着用fgets读取剩余内容,会把文件指针直接移到文件末尾,后续fscanf无法再读取其他单词。
  • 缺少单词比对逻辑:代码里没有实现单词匹配和计数的核心逻辑,只是单纯打印读取到的第一个单词。
  • 冗余的文件名复制:没必要手动复制文件名到f数组,直接使用传入的file参数即可。
  • 未处理标点符号:输入中的want.会被当成和want不同的单词,导致统计错误。
  • 无重复统计防护:即使逻辑正确,也会重复统计同一个单词多次。

修复后的代码
#include <stdio.h>
#include <string.h>
#include <ctype.h>

#define MAX_WORDS 100  // 假设最多统计100个不同单词
#define MAX_WORD_LEN 100

// 存储单词和计数的结构体
typedef struct {
    char word[MAX_WORD_LEN];
    int count;
} WordCount;

void CountAllOccurren(char *file) {
    WordCount wordList[MAX_WORDS] = {0};  // 初始化单词列表
    int wordCount = 0;  // 已统计的不同单词数量
    char currentWord[MAX_WORD_LEN];
    FILE* file_ptr = fopen(file, "r");

    if (!file_ptr) {
        printf("无法打开文件\n");
        return;
    }

    // 逐个读取单词
    while (fscanf(file_ptr, "%s", currentWord) != EOF) {
        // 去除单词末尾的标点符号
        char *end = currentWord + strlen(currentWord) - 1;
        while (end >= currentWord && ispunct((unsigned char)*end)) {
            *end = '\0';
            end--;
        }

        // 跳过空字符串(比如如果单词全是标点的情况)
        if (strlen(currentWord) == 0) {
            continue;
        }

        // 检查是否已经统计过该单词
        int found = 0;
        for (int i = 0; i < wordCount; i++) {
            if (strcmp(wordList[i].word, currentWord) == 0) {
                wordList[i].count++;
                found = 1;
                break;
            }
        }

        // 没找到就添加新单词
        if (!found && wordCount < MAX_WORDS) {
            strcpy(wordList[wordCount].word, currentWord);
            wordList[wordCount].count = 1;
            wordCount++;
        }
    }

    fclose(file_ptr);

    // 输出统计结果
    for (int i = 0; i < wordCount; i++) {
        if (i > 0) {
            printf(" - ");
        }
        printf("%s:%d", wordList[i].word, wordList[i].count);
    }
    printf("\n");
}

// 测试用例
int main() {
    CountAllOccurren("test.txt");
    return 0;
}

关键说明
  • 结构体存储:用WordCount结构体保存每个单词及其出现次数,避免重复统计。
  • 标点处理:通过ispunct函数识别并去除单词末尾的标点,确保want和want.被视为同一个单词。
  • 重复检查:每次读取单词后,遍历已统计的单词列表,找到匹配项就增加计数,否则添加新条目。
  • 错误处理:增加了文件打开失败的判断,避免程序崩溃。

测试输入文件test.txt内容:

i want to help to want want.

运行后输出:

i:1 - want:3 - to:2 - help:1

内容的提问来源于stack exchange,提问作者quoc nguyen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 10:35:29