如何用C语言读取文件并统计不同单词数量(补全现有代码)
补全代码实现统计不同单词数
看起来你已经搞定了总单词数的统计,现在要实现去重后的不同单词计数对吧?我给你补全代码,用链表来存储唯一单词——这是处理不确定数量元素的灵活方案,不用担心数组大小不够的问题。
完整代码(含补全部分)
#include <stdio.h> #include <stdlib.h> #include <string.h> #include <ctype.h> // 用于处理大小写统一(可选) #define MAX 80 typedef char string[MAX+1]; // 定义链表节点,存储唯一单词 typedef struct Node { string word; struct Node *next; } Node; // 检查单词是否已存在于链表中 int wordExists(Node *head, const string target) { Node *current = head; while (current != NULL) { // 若需忽略大小写,可替换为strcasecmp(Linux/macOS)或_stricmp(Windows) if (strcmp(current->word, target) == 0) { return 1; // 存在返回1 } current = current->next; } return 0; // 不存在返回0 } // 向链表头部添加新单词 Node* addWord(Node *head, const string newWord) { Node *newNode = (Node*)malloc(sizeof(Node)); if (newNode == NULL) { printf("内存分配失败!\n"); exit(1); } strcpy(newNode->word, newWord); newNode->next = head; return newNode; } // 统计链表中唯一单词的数量 int countUniqueWords(Node *head) { int count = 0; Node *current = head; while (current != NULL) { count++; current = current->next; } return count; } // 释放链表内存,避免泄漏 void freeList(Node *head) { Node *temp; while (head != NULL) { temp = head; head = head->next; free(temp); } } void main() { FILE *fp; string filename, word; int totalWords = 0; Node *uniqueWordsHead = NULL; // 唯一单词链表的头节点 printf("请输入要读取的文件名:"); scanf("%s", filename); fp = fopen(filename, "r"); if (fp == NULL) { printf("无法打开文件 %s\n", filename); exit(1); } // 读取文件中的每个单词 while (fscanf(fp, "%s", word) != EOF) { totalWords++; // 统计总单词数 // 可选:统一转换为小写,让Hello和hello被视为同一个单词 // for (int i = 0; word[i]; i++) { // word[i] = tolower(word[i]); // } // 若单词未出现过,添加到唯一单词链表 if (!wordExists(uniqueWordsHead, word)) { uniqueWordsHead = addWord(uniqueWordsHead, word); } } fclose(fp); // 计算并输出结果 int uniqueCount = countUniqueWords(uniqueWordsHead); printf("总单词数:%d\n", totalWords); printf("不同单词数:%d\n", uniqueCount); // 清理链表内存 freeList(uniqueWordsHead); }
补全部分的核心逻辑解析
- 链表结构设计:用链表存储已经出现过的单词,每个节点保存一个单词和下一个节点的指针,动态扩展不会浪费内存。
- 唯一性检查:每次读取新单词后,遍历链表对比,确保只有没出现过的单词才会被加入链表。
- 内存管理:最后释放链表的所有节点,避免C语言常见的内存泄漏问题。
- 可选优化:代码里注释了大小写统一的逻辑,如果你需要忽略大小写差异,取消注释即可。
这个方案逻辑清晰,和你原有的总单词统计逻辑完美整合,运行后就能同时输出总单词数和去重后的不同单词数啦。
内容的提问来源于stack exchange,提问作者J'Neal
相关产品推荐
相关产品推荐

