统计文本文件单词出现次数的C程序异常,如何修复strtok问题?
问题修复方案
你的代码出现异常输出,核心问题出在strtok的使用和几个逻辑细节上,以下是具体修复点:
1. 补全strtok的分隔符集合
你只把空格作为分隔符,但输入文本里有逗号、句号、括号、方括号、换行符等非单词字符,这些会被粘在单词上(比如computation,、information.[1]),导致相同单词被误判为不同项,甚至出现乱码。
修复时把分隔符改成包含所有非单词字符的集合:
// 第一次调用strtok char *word = strtok(line, " \n,.()[]"); // 后续的strtok也要用同样的分隔符 word = strtok(NULL, " \n,.()[]");
2. 处理大小写差异(可选但实用)
输入里的Computer和computer会被当成不同单词,如果需要统一统计,可以新增一个转小写的函数,在处理单词时统一格式:
char *to_lower(char *s) { char *p = s; while (*p) { *p = tolower((unsigned char)*p); p++; } return s; }
然后在获取到word后立刻调用:
to_lower(word);
3. 避免数组越界
你初始化的数组大小是1024,如果文本里的不同单词超过这个数量,会导致内存越界,产生无意义输出。可以在每次新增单词前检查是否需要扩容:
if (i >= size) { size *= 2; words = realloc(words, size * sizeof(char*)); wordcount = realloc(wordcount, size * sizeof(int)); // 初始化新分配的内存,避免脏数据 memset(words + i, 0, (size - i) * sizeof(char*)); memset(wordcount + i, 0, (size - i) * sizeof(int)); }
4. 优化输出循环
原来的输出循环遍历了整个1024大小的数组,会输出空元素。改成遍历实际统计的单词数量i即可:
for (int idx = 0; idx < i; idx++) { if (words[idx] != NULL) { printf("%s %d\n", words[idx], wordcount[idx]); } }
修复后的完整代码
#include <stdio.h> #include <string.h> #include <stdlib.h> #include <ctype.h> char *to_lower(char *s) { char *p = s; while (*p) { *p = tolower((unsigned char)*p); p++; } return s; } int main(int argc, char *argv[]){ FILE *file=fopen(argv[1],"r"); if(!file) { perror("File opening failed"); return 1; } int size = 1024; int *wordcount = calloc(size, sizeof(int)); char **words = calloc(size, sizeof(char*)); char line[1024]; int p=0; int i=0; while (fgets(line, sizeof(line), file) != NULL) { char *word = strtok(line, " \n,.()[]"); while (word != NULL) { to_lower(word); p = 0; for (int j=0;j<i;j++){ if (strcmp(words[j], word) == 0){ wordcount[j]++; p = 1; break; } } if (!p){ if (i >= size) { size *= 2; words = realloc(words, size * sizeof(char*)); wordcount = realloc(wordcount, size * sizeof(int)); if (!words || !wordcount) { perror("Realloc failed"); return 1; } memset(words + i, 0, (size - i) * sizeof(char*)); memset(wordcount + i, 0, (size - i) * sizeof(int)); } words[i] = strdup(word); wordcount[i]++; i++; } word = strtok(NULL, " \n,.()[]"); } } for (int idx = 0; idx < i; idx++) { if (words[idx] != NULL) { printf("%s %d\n", words[idx], wordcount[idx]); } } fclose(file); for (int idx = 0; idx < i; idx++) { if (words[idx] != NULL) { free(words[idx]); } } free(words); free(wordcount); return 0; }
内容的提问来源于stack exchange,提问作者Giorgos Zigouris
相关产品推荐
相关产品推荐

