如何统计文本文件每行的单词数?C语言开发相关问询
每行单词数统计的C语言实现
如果你的程序现在没法区分每行的单词,核心问题是数据结构没按行存储,或者读取文件时没做行级的单词分割统计。下面给你两种可行的实现方案:
1. 推荐方案:按行存储数据并统计
先调整你的Wordle结构,让它能存每行的单词列表和对应的单词数,这样后续处理会非常方便:
// 存储单行的单词和数量 typedef struct { char **words; // 该行的所有单词 int word_count; // 该行的单词总数 } LineData; // 存储整个文件的行数据 typedef struct { LineData *lines; // 所有行的数组 int line_count; // 文件总行数 } Wordle;
读取文件时统计每行单词数
用getline逐行读取文件,再用strtok分割每行的单词,同时计数:
#include <stdio.h> #include <stdlib.h> #include <string.h> #include <ctype.h> // 分割一行字符串为单词,返回单词数 int split_line(char *line, char ***words) { int count = 0; // 按空格、制表符、换行符分割单词 char *token = strtok(line, " \t\n"); while (token != NULL) { // 可选:去掉单词末尾的标点(比如示例里的"good?"变成"good") size_t len = strlen(token); if (ispunct(token[len-1])) { token[len-1] = '\0'; } // 动态分配内存存单词 *words = realloc(*words, (count + 1) * sizeof(char*)); (*words)[count] = malloc(strlen(token) + 1); strcpy((*words)[count], token); count++; token = strtok(NULL, " \t\n"); } return count; } // 读取文件并填充Wordle结构 Wordle* load_file(const char *filename) { Wordle *wordle = malloc(sizeof(Wordle)); wordle->line_count = 0; wordle->lines = NULL; FILE *fp = fopen(filename, "r"); if (!fp) { perror("打开文件失败"); free(wordle); return NULL; } char *line = NULL; size_t line_cap = 0; ssize_t line_len; // 逐行读取 while ((line_len = getline(&line, &line_cap, fp)) != -1) { // 跳过空行 if (line_len <= 1) continue; char **words = NULL; int word_num = split_line(line, &words); // 扩展行数组 wordle->lines = realloc(wordle->lines, (wordle->line_count + 1) * sizeof(LineData)); wordle->lines[wordle->line_count].words = words; wordle->lines[wordle->line_count].word_count = word_num; wordle->line_count++; } free(line); fclose(fp); return wordle; }
根据单词数控制输出
现在打印时可以直接拿到每行的单词数,按需调整格式:
void print_content(Wordle *wordle) { for (int i = 0; i < wordle->line_count; i++) { LineData *curr_line = &wordle->lines[i]; printf("第%d行,单词数:%d\n", i+1, curr_line->word_count); // 示例:根据单词数调整格式 if (curr_line->word_count == 5) { // 5个单词时居中输出 printf(" "); for (int j = 0; j < curr_line->word_count; j++) { printf("%s ", curr_line->words[j]); } } else if (curr_line->word_count == 3 || curr_line->word_count ==4) { // 3/4个单词时左对齐 for (int j = 0; j < curr_line->word_count; j++) { printf("%s ", curr_line->words[j]); } } printf("\n"); } }
别忘了释放内存
用完后要清理分配的内存,避免泄漏:
void free_wordle(Wordle *wordle) { for (int i = 0; i < wordle->line_count; i++) { for (int j = 0; j < wordle->lines[i].word_count; j++) { free(wordle->lines[i].words[j]); } free(wordle->lines[i].words); } free(wordle->lines); free(wordle); }
2. 临时方案:不修改现有结构
如果你不想改现有的allwords一维数组,那得提前在读取文件时记录每行的单词数,存在另一个数组里,比如:
// 假设你已经有line_word_counts数组,存了每行的单词数 int line_word_counts[] = {4, 2, 5}; // 对应示例文本的三行 int curr_word = 0; for (int i = 0; i < wordle->leng; i++) { int count = line_word_counts[i]; printf("第%d行单词数:%d\n", i+1, count); // 输出该行的所有单词 for (int j = 0; j < count; j++) { printf("%s ", wordle->allwords[curr_word++]); } printf("\n"); }
这种方式局限性大,一旦文件内容变化就得手动更新计数数组,所以还是推荐第一种按行存储的方案。
内容的提问来源于stack exchange,提问作者dakota kellogg
相关产品推荐
相关产品推荐

