You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C语言新手求助:分词存储结构体及词频统计排序实现困惑

Hey there! Totally get it when you're starting out with C and existing answers don't quite click for your own code—let's break this down step by step, focusing exactly on the struct-based approach you want to use. This is a super common problem, and structs are actually the perfect tool here.

1. First, Design Your Word Count Struct

You need a way to tie each unique word to its frequency. A struct with two fields will do the trick:

#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <ctype.h>

// Define the struct to hold a word and its count
typedef struct {
    char *word;  // Stores the actual word (dynamic string)
    int count;   // Number of times the word appears
} WordEntry;

We use a char* instead of a fixed-size array because words can vary in length—this lets us allocate exactly the space we need for each word.

2. Use a Dynamic Array of Structs

Since you don't know how many unique words you'll have upfront, a dynamic array (resizable) of WordEntry is the way to go. We'll track three things:

  • The array itself (word_list)
  • How many elements the array can currently hold (list_size)
  • How many unique words we've actually added (word_count)

Here's how to initialize these:

WordEntry *word_list = NULL;
int list_size = 0;       // Starts empty, we'll resize as needed
int word_count = 0;      // Tracks actual unique words

Adding Words to the Array

Next, write a helper function to check if a word already exists in the array. If it does, increment its count. If not, add a new WordEntry (and resize the array if it's full):

void add_word(WordEntry **list, int *list_size, int *word_count, const char *word) {
    // First, check if the word is already in our list
    for (int i = 0; i < *word_count; i++) {
        if (strcmp((*list)[i].word, word) == 0) {
            (*list)[i].count++;
            return; // No need to add a new entry
        }
    }

    // If array is full, resize it (we'll double the size each time for efficiency)
    if (*word_count >= *list_size) {
        int new_size = (*list_size == 0) ? 4 : *list_size * 2;
        WordEntry *temp = realloc(*list, new_size * sizeof(WordEntry));
        if (temp == NULL) {
            perror("Failed to resize word list");
            exit(EXIT_FAILURE);
        }
        *list = temp;
        *list_size = new_size;
    }

    // Allocate space for the new word and copy it in
    (*list)[*word_count].word = malloc(strlen(word) + 1); // +1 for null terminator
    if ((*list)[*word_count].word == NULL) {
        perror("Failed to allocate space for word");
        exit(EXIT_FAILURE);
    }
    strcpy((*list)[*word_count].word, word);
    (*list)[*word_count].count = 1;
    (*word_count)++;
}

Note the double pointers (**list)—we need these because we're modifying the original array pointer when we realloc.

3. Clean Up Words (Optional But Important)

Before adding a word to your list, you'll want to normalize it: convert to lowercase (so "Hello" and "hello" count as the same word) and strip trailing punctuation (like commas or periods). Here's a quick helper for that:

void clean_word(char *word) {
    // Convert all characters to lowercase
    for (int i = 0; word[i] != '\0'; i++) {
        word[i] = tolower(word[i]);
    }

    // Remove trailing punctuation
    int len = strlen(word);
    while (len > 0 && ispunct(word[len - 1])) {
        len--;
    }
    word[len] = '\0'; // Null-terminate the cleaned word
}
4. Sort by Frequency (Descending)

Once you've collected all words, use C's built-in qsort function to sort the array. You just need to write a comparison function that sorts first by frequency (highest to lowest), and by alphabetical order if frequencies are equal:

int compare_word_entries(const void *a, const void *b) {
    const WordEntry *entry_a = (const WordEntry *)a;
    const WordEntry *entry_b = (const WordEntry *)b;

    // First sort by count descending
    if (entry_b->count != entry_a->count) {
        return entry_b->count - entry_a->count;
    }
    // If counts are equal, sort alphabetically ascending
    return strcmp(entry_a->word, entry_b->word);
}

Then call it like this:

qsort(word_list, word_count, sizeof(WordEntry), compare_word_entries);
5. Put It All Together (Example Workflow)

Here's a quick snippet showing how to tie this all together with reading from stdin:

int main() {
    char buffer[256]; // Buffer to hold each word we read

    // Read words from stdin (you can adapt this to read from a file with fopen/fscanf)
    while (scanf("%s", buffer) != EOF) {
        clean_word(buffer);
        // Skip empty strings (in case we stripped all punctuation from a word like "!!!")
        if (strlen(buffer) > 0) {
            add_word(&word_list, &list_size, &word_count, buffer);
        }
    }

    // Sort the word list
    qsort(word_list, word_count, sizeof(WordEntry), compare_word_entries);

    // Print the results
    printf("Word Frequency (Highest to Lowest):\n");
    for (int i = 0; i < word_count; i++) {
        printf("%-15s %d\n", word_list[i].word, word_list[i].count);
    }

    // Clean up memory (don't forget this!)
    for (int i = 0; i < word_count; i++) {
        free(word_list[i].word);
    }
    free(word_list);

    return 0;
}
Key Tips for a C Newbie
  • Always check the return values of malloc and realloc—they can fail if you run out of memory.
  • Don't forget to free all allocated memory to avoid leaks.
  • If you're reading from a file instead of stdin, replace scanf with fscanf using a file pointer from fopen.
  • The strtok function can be useful if you need to split text on spaces and other whitespace, but scanf("%s") works for basic word splitting (it stops at whitespace).

内容的提问来源于stack exchange,提问作者E.Bille

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:29:49