C语言统计文本文件所有单词出现次数的代码问题求助
问题分析与修复
你的代码存在几个关键问题导致无法正确统计单词出现次数:
- 文件读取逻辑错误:第一次用
fscanf读取一个单词后,紧接着用fgets读取剩余内容,会把文件指针直接移到文件末尾,后续fscanf无法再读取其他单词。 - 缺少单词比对逻辑:代码里没有实现单词匹配和计数的核心逻辑,只是单纯打印读取到的第一个单词。
- 冗余的文件名复制:没必要手动复制文件名到
f数组,直接使用传入的file参数即可。 - 未处理标点符号:输入中的
want.会被当成和want不同的单词,导致统计错误。 - 无重复统计防护:即使逻辑正确,也会重复统计同一个单词多次。
修复后的代码
#include <stdio.h> #include <string.h> #include <ctype.h> #define MAX_WORDS 100 // 假设最多统计100个不同单词 #define MAX_WORD_LEN 100 // 存储单词和计数的结构体 typedef struct { char word[MAX_WORD_LEN]; int count; } WordCount; void CountAllOccurren(char *file) { WordCount wordList[MAX_WORDS] = {0}; // 初始化单词列表 int wordCount = 0; // 已统计的不同单词数量 char currentWord[MAX_WORD_LEN]; FILE* file_ptr = fopen(file, "r"); if (!file_ptr) { printf("无法打开文件\n"); return; } // 逐个读取单词 while (fscanf(file_ptr, "%s", currentWord) != EOF) { // 去除单词末尾的标点符号 char *end = currentWord + strlen(currentWord) - 1; while (end >= currentWord && ispunct((unsigned char)*end)) { *end = '\0'; end--; } // 跳过空字符串(比如如果单词全是标点的情况) if (strlen(currentWord) == 0) { continue; } // 检查是否已经统计过该单词 int found = 0; for (int i = 0; i < wordCount; i++) { if (strcmp(wordList[i].word, currentWord) == 0) { wordList[i].count++; found = 1; break; } } // 没找到就添加新单词 if (!found && wordCount < MAX_WORDS) { strcpy(wordList[wordCount].word, currentWord); wordList[wordCount].count = 1; wordCount++; } } fclose(file_ptr); // 输出统计结果 for (int i = 0; i < wordCount; i++) { if (i > 0) { printf(" - "); } printf("%s:%d", wordList[i].word, wordList[i].count); } printf("\n"); } // 测试用例 int main() { CountAllOccurren("test.txt"); return 0; }
关键说明
- 结构体存储:用
WordCount结构体保存每个单词及其出现次数,避免重复统计。 - 标点处理:通过
ispunct函数识别并去除单词末尾的标点,确保want和want.被视为同一个单词。 - 重复检查:每次读取单词后,遍历已统计的单词列表,找到匹配项就增加计数,否则添加新条目。
- 错误处理:增加了文件打开失败的判断,避免程序崩溃。
测试输入文件test.txt内容:
i want to help to want want.
运行后输出:
i:1 - want:3 - to:2 - help:1
内容的提问来源于stack exchange,提问作者quoc nguyen
相关产品推荐
相关产品推荐

