You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C语言split_words函数内存分配异常及越界问题求助

问题排查:字符串分割函数返回异常与数组越界问题

任务要求实现split_words函数,将输入文本分割为单词并存入char**数组,数组末尾需带NULL指针;无单词时返回指向NULL的指针。

当前遇到的问题:

  • 本地CLion测试正常,但课程测试程序报错,称函数返回NULL而非已分配内存的指针
  • 测试提示设置数组末尾*(result + next_word) = NULL时存在数组越界访问

函数实现代码

split_words函数

char** split_words(const char *text){
    if (text == NULL) return NULL;

    //fun() to count words in given text
    int words_number = count_words(text);

    //allocate memory for result array, size: words_number * pointer for each word
    char **result = (char**)malloc((words_number + 1) * sizeof(char*));
    if (result == NULL){
        return NULL;
    }

    //temporary word variable to save wrods in result array
    char *word = (char*)malloc(50 * sizeof(char));
    if (word == NULL){
        destroy(result);
        return NULL;
    }

    int i = 0, next_word=0;
    while (i < (int)strlen(text)){
        while (i < (int)strlen(text)){
            if ((*(text + i) >= 'a' && *(text + i) <= 'z') || (*(text + i) >= 'A' && *(text + i) <= 'Z')) break;
            i++;
        }
        int j = 0;
        while (i < (int)strlen(text)){
            if (strchr(" '.,{}[]/\\&quot;-;:&quot;", *(text + i)) != NULL) break;
            *(word + j) = *(text + i);
            i++;
            j++;
        }
        if (j > 0){
            // add
            *(word + j) = '\0';
            // allocate memory for each word in result array
            *(result + next_word) = (char*)malloc((strlen(word) + 1) * sizeof(char));
            if(*(result + next_word) == NULL){
//                destroy(result);
                free(word);
                return NULL;
            }
            strcpy(*(result + next_word), word);
            next_word++;
        }
    }

    // add NULL pointer at the end of words array
    *(result + next_word) = NULL;
    free(word);
    return result;
}

destroy函数

void destroy(char **words){
    if (words == NULL) return;

    int i = 0;
    while (*(words + i) != NULL){
        free(*(words + i));
        i++;
    }

    free(words);
}

count_words函数

int count_words(const char *text){
    if (text == NULL) return 0;
    int counter = 0, state = 1;

    for (int i=0 ; *(text + i) != '\0'; i++){
        if (strchr(" '.,{}[]/\\&quot;-;:&quot;", *(text + i)) != NULL){
            state = 1;
        } else if (state == 1){
            state = 0;
            counter++;
        }
    }

    return counter;
}

问题排查方向

  1. 单词判定逻辑不一致
    count_words与split_words的分隔符判定逻辑存在隐性差异:split_words的第一个循环仅跳过非字母字符,而count_words只要字符不在分隔符列表就判定为单词起始。若输入包含数字、下划线等非字母但非分隔符的字符,会导致count_words统计的单词数与实际分割出的数量不匹配,进而引发result数组的内存分配大小错误,最终导致越界访问。

  2. 分隔符字符串转义错误
    代码中的分隔符字符串包含\&quot;这类HTML转义字符,而非C语言的正确转义格式(\"),这会导致strchr的匹配逻辑完全错误,使得count_words统计的单词数严重偏离实际值,直接引发result数组的内存分配大小与实际需求不匹配。

  3. 临时缓冲区长度限制
    临时word缓冲区固定分配50字节,若遇到长度超过49的单词,会出现缓冲区写越界,破坏内存中的next_word或result指针值,进而引发后续的越界访问或返回错误指针。

  4. 内存分配失败时的资源泄漏
    当某个单词的malloc失败时,仅释放了临时word缓冲区,未释放已分配的单词指针和result数组本身,不仅会造成内存泄漏,还可能破坏测试环境的内存状态,引发异常报错。

  5. 重复调用strlen的潜在风险
    每次循环都调用strlen(text),虽不会直接引发错误,但会降低效率;若text在运行中被意外修改(尽管参数是const),可能导致循环条件失效。

修复建议

  1. 统一单词判定逻辑
    提取辅助函数is_delimiter(char c),让count_words和split_words共用同一套分隔符判定逻辑,避免统计与分割的数量差异:

    static int is_delimiter(char c) {
        return strchr(" '.,{}[]/\\\"-;:\"", c) != NULL;
    }
    
  2. 修复分隔符字符串转义
    将分隔符字符串修正为C语言正确格式:" '.,{}[]/\\\"-;:\"",确保strchr能正确匹配分隔符。

  3. 动态处理单词缓冲区
    移除固定大小的临时缓冲区,改为先定位单词的起始和结束位置,直接分配对应大小的内存存储单词,避免缓冲区越界:

    // 替换原临时word相关逻辑
    while (i < text_len) {
        // 跳过分隔符
        while (i < text_len && is_delimiter(text[i])) i++;
        if (i >= text_len) break;
        // 定位单词结束位置
        int start = i;
        while (i < text_len && !is_delimiter(text[i])) i++;
        int word_len = i - start;
        // 分配单词内存
        result[next_word] = malloc(word_len + 1);
        if (!result[next_word]) {
            // 清理已分配资源
            for (int k=0; k<next_word; k++) free(result[k]);
            free(result);
            return NULL;
        }
        strncpy(result[next_word], text + start, word_len);
        result[next_word][word_len] = '\0';
        next_word++;
    }
    
  4. 完善内存分配失败的清理逻辑
    当单词malloc失败时,先释放已分配的所有单词指针,再释放result数组,避免内存泄漏和环境破坏。

  5. 提前计算字符串长度
    提前调用一次strlen(text)并保存结果,避免循环中重复调用,提升效率并避免潜在风险。

内容的提问来源于stack exchange,提问作者anocyney_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 17:05:32