使用fgets与strcmp统计文件字符串出现次数时匹配异常问题
统计文件中目标字符串出现次数的问题解决
问题描述
需要实现一个函数,接收FILE指针和字符串指针,区分大小写统计文件中该字符串的出现次数并返回整数。测试时目标字符串为"dog",test.txt中实际存在3次"dog",但程序始终返回0。
初始尝试与问题
最初使用fgets读取每行内容,截断换行符后用strcmp匹配,代码如下:
#include <stdio.h> #include <string.h> int searchAndCount(FILE *fp, const char *searchWord) { int count = 0; char buffer[4096]; while ((fgets(buffer, sizeof(buffer), fp)) != NULL) { buffer[strlen(buffer) - 1] = '\0'; if (strcmp(buffer, searchWord) == 0) { count++; } } return count; } int main() { FILE *fp; int searchedWord; const char *c = "dog"; fp = fopen("test.txt", "r"); if (fp == NULL) { perror("File couldn't open properly"); return 1; } searchedWord = searchAndCount(fp, c); printf("The word 'dog' occurs %d-times in the file\n", searchedWord); fclose(fp); return 0; }
test.txt内容:
dog dog dogoggo dog.
运行结果:
The word 'dog' occurs 0-times in the file
问题核心:fgets读取整行内容,strcmp要求整行完全等于"dog"才会计数,但测试文件中一行包含多个单词,因此匹配失败,返回0。
第一次改进与遗留问题
根据提示研究strtok后,改用fscanf按单词读取,代码如下:
int searchAndCount(FILE *fp, const char *searchWord) { int count = 0; char buffer[4096]; while ((fscanf(fp, "%4095s", buffer)) == 1) { if (strcmp(buffer, searchWord) == 0) { count++; } } return count; }
该方案解决了多行多单词的匹配问题,但无法统计带标点的"dog."——因为%s会把"dog."作为一个完整单词,和"dog"不匹配。
最终解决方案
使用fscanf的格式控制,先跳过非字母字符,再读取连续的字母作为单词,剥离标点后准确匹配目标字符串:
int searchAndCount(FILE *fp, const char *searchWord) { int count = 0; char buffer[4096]; for (;;) { fscanf(fp, "%*[^a-zA-Z]"); if (fscanf(fp, "%4095[a-zA-Z]", buffer) != 1) break; if (strcmp(buffer, searchWord) == 0) count++; } return count; }
%*[^a-zA-Z]:跳过所有非字母字符,*表示读取但不存储结果%4095[a-zA-Z]:读取最多4095个连续字母,存入buffer- 循环直到无法读取到有效字母单词,每次读取后与目标字符串匹配,匹配成功则计数加1
内容的提问来源于stack exchange,提问作者R0gue
相关产品推荐
相关产品推荐

