You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用strtok实现单词生成器遇段错误,如何处理不确定单词数?

C语言单词生成器段错误分析与动态扩容解决方案

段错误的核心原因

你代码里的words数组是用固定大小NUM_WORDS分配的,当文本中的实际单词数量超过这个值时,words[num_words]会越界访问非法内存,直接触发Segmentation Fault。这是处理不确定数量元素时的典型问题——固定容器无法适配可变长度的输入。

解决思路与修正代码

要处理数量不确定的单词,必须用动态扩容的方式管理存储单词的指针数组,同时优化单词处理逻辑避免其他潜在问题:

关键优化点

  • 动态扩容数组:初始分配一个合理的小容量,当元素数量达到容量上限时,用realloc扩容(通常每次翻倍,减少扩容次数)。
  • 使用线程安全的strtok_r替代strtok:strtok依赖全局静态变量,多线程或嵌套调用时会出问题,strtok_r更可靠。
  • 复制原文本:strtok系列函数会修改输入字符串,复制一份避免破坏原始数据。
  • 精确复制单词:根据处理后的长度复制内容,避免保留多余空白字符,同时跳过空单词(比如连续空白产生的无效内容)。

修正后的完整代码

#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <ctype.h>

char** split_words(const char* text, int* out_num_words) {
    if (text == NULL || out_num_words == NULL) {
        return NULL;
    }

    int capacity = 8; // 初始容量,可按需调整
    char** words = malloc(sizeof(char*) * capacity);
    if (words == NULL) {
        return NULL;
    }

    int num_words = 0;
    // 复制原文本,防止strtok_r修改原始数据
    char* text_copy = strdup(text);
    if (text_copy == NULL) {
        free(words);
        return NULL;
    }

    char* save_ptr;
    // 用所有空白字符作为分隔符,适配空格、制表符、换行等
    char* word = strtok_r(text_copy, " \t\n\r\f\v", &save_ptr);
    while (word != NULL) {
        // 去掉单词尾部的空白字符
        int len = strlen(word);
        while (len > 0 && isspace((unsigned char)word[len - 1])) {
            len--;
        }

        // 跳过空单词(比如连续空白产生的无效内容)
        if (len == 0) {
            word = strtok_r(NULL, " \t\n\r\f\v", &save_ptr);
            continue;
        }

        // 检查是否需要扩容
        if (num_words >= capacity) {
            capacity *= 2;
            char** temp = realloc(words, sizeof(char*) * capacity);
            if (temp == NULL) {
                // 内存分配失败,清理已分配资源
                for (int i = 0; i < num_words; i++) {
                    free(words[i]);
                }
                free(words);
                free(text_copy);
                return NULL;
            }
            words = temp;
        }

        // 为处理后的单词分配内存并复制
        words[num_words] = malloc(len + 1);
        if (words[num_words] == NULL) {
            // 清理资源
            for (int i = 0; i < num_words; i++) {
                free(words[i]);
            }
            free(words);
            free(text_copy);
            return NULL;
        }
        strncpy(words[num_words], word, len);
        words[num_words][len] = '\0'; // 手动添加字符串结束符

        num_words++;
        word = strtok_r(NULL, " \t\n\r\f\v", &save_ptr);
    }

    free(text_copy);
    // 可选:缩容到实际需要的大小,节省内存
    char** final_words = realloc(words, sizeof(char*) * num_words);
    if (final_words != NULL) {
        words = final_words;
    }

    *out_num_words = num_words;
    return words;
}

// 使用示例
int main() {
    const char* text = "some text without comma   and   extra spaces\twith tab\nnew line word";
    int num_words;
    char** words = split_words(text, &num_words);
    if (words == NULL) {
        fprintf(stderr, "Failed to split words\n");
        return 1;
    }

    for (int i = 0; i < num_words; i++) {
        printf("Word %d: %s\n", i+1, words[i]);
        free(words[i]);
    }
    free(words);
    return 0;
}

使用说明

  • 调用split_words时传入文本和一个用于接收单词数量的指针,函数返回存储单词的指针数组。
  • 使用完毕后要逐个释放每个单词的内存,最后释放数组本身,避免内存泄漏。

内容的提问来源于stack exchange,提问作者kenjsch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 05:45:24