如何在C语言中通过逐字符读取TXT文件将每个单词存储到二维数组中
Problem Statement
I've written the following C code to count characters, spaces, and words in a text file:
int calc_stats(void) { char in_name[80]; FILE* in_file; int ch, character = 0, space = 0, words = 0; char str[30]; int i; printf("Enter file name:\n"); scanf("%s", in_name); in_file = fopen(in_name, "r"); if (in_file == NULL) printf("Can't open %s for reading.\n", in_name); else { while ((ch = fgetc(in_file)) != EOF) { character++; if (ch == ' ') { space++; } if (ch == ' ' || ch == '\t' || ch == '\n' || ch == '\0') { words++; strcat(str, " "); } else { strcat(str, ch); } } fclose(in_file); printf("\nNumber of characters = %d", character); printf("\nNumber of characters without space = %d", character - space); printf("\nNumber of words = %d", words); } return 0; }My goal is to store each recognized word into a 2D array while reading the file. However, right now I'm reading character by character with
ch = fgetc(in_file)and judging characters, and I don't know how to combine characters into words and store them in a 2D array. I hope to get relevant technical help.
Solution
Hey there! Let's fix your code step by step to store words into a 2D array, while also cleaning up some existing issues in your original implementation.
First, let's address the problems in your current code:
- The
strarray isn't initialized, sostrcatwill lead to undefined behavior. strcat(str, ch)is invalid becausestrcatexpects a C-style string (achar*), not a singlechar.- Your word counting logic is off: it increments
wordsevery time it hits a separator, which means multiple consecutive spaces/tabs/newlines will count as multiple words (which isn't correct).
Now, here's how to implement the word storage into a 2D array:
- Define a 2D array to hold words: We'll set a reasonable limit (e.g., 100 words max, each up to 29 characters plus the null terminator).
- Track current state: Use two indices—one for the current word we're building, another for the current character position within that word.
- Build words character by character: When we read a non-separator character, add it to the current word. When we hit a separator, finalize the current word (if it's not empty) and move to the next one.
- Handle the final word: After the loop ends, don't forget to save the last word if it wasn't terminated by a separator.
Here's the revised code with these changes:
#include <stdio.h> #include <string.h> int calc_stats(void) { char in_name[80]; FILE* in_file; int ch, character_count = 0, space_count = 0, word_count = 0; // Define a 2D array to store words: 100 words, each up to 29 chars + null terminator char words_array[100][30] = {0}; // Initialize to zero to avoid garbage values int current_word_idx = 0; // Index for the current word in the array int current_char_idx = 0; // Index for the current character in the current word printf("Enter file name:\n"); scanf("%s", in_name); in_file = fopen(in_name, "r"); if (in_file == NULL) { printf("Can't open %s for reading.\n", in_name); return 1; // Return an error code instead of continuing } while ((ch = fgetc(in_file)) != EOF) { character_count++; // Check if the character is a separator (space, tab, newline) if (ch == ' ' || ch == '\t' || ch == '\n') { space_count++; // If we were building a word, finalize it if (current_char_idx > 0) { words_array[current_word_idx][current_char_idx] = '\0'; // Add null terminator current_word_idx++; word_count++; current_char_idx = 0; // Reset for next word } } else { // Add the character to the current word (if we haven't hit the word length limit) if (current_char_idx < 29) { // Leave space for the null terminator words_array[current_word_idx][current_char_idx++] = ch; } } } // Handle the last word in case the file doesn't end with a separator if (current_char_idx > 0) { words_array[current_word_idx][current_char_idx] = '\0'; word_count++; current_word_idx++; } fclose(in_file); // Print the stats printf("\nNumber of characters = %d", character_count); printf("\nNumber of characters without space = %d", character_count - space_count); printf("\nNumber of words = %d", word_count); // Optional: Print all stored words to verify printf("\n\nStored words:\n"); for (int i = 0; i < word_count; i++) { printf("- %s\n", words_array[i]); } return 0; } int main() { calc_stats(); return 0; }
Key points explained:
- Initialization: The
words_arrayis initialized to all zeros, so each word starts with a clean slate, and we don't have to worry about leftover garbage characters. - Word building: For every non-separator character, we add it to the current position in the 2D array. We cap the character index at 29 to prevent buffer overflow (since each word array is 30 elements long, one for the null terminator).
- Separator handling: When we hit a separator, we only finalize the word if we were actually building one (i.e.,
current_char_idx > 0). This avoids counting empty "words" from multiple consecutive separators. - Final word check: After the loop ends, if
current_char_idxis greater than 0, it means we were in the middle of a word when we reached EOF, so we need to save that word too.
This implementation will correctly store each word into the 2D array while maintaining accurate counts for characters, spaces, and words.
内容的提问来源于stack exchange,提问作者Manolis Iordanidis

