C语言split_words函数内存分配异常及越界问题求助
任务要求实现split_words函数,将输入文本分割为单词并存入char**数组,数组末尾需带NULL指针;无单词时返回指向NULL的指针。
当前遇到的问题:
- 本地CLion测试正常,但课程测试程序报错,称函数返回NULL而非已分配内存的指针
- 测试提示设置数组末尾
*(result + next_word) = NULL时存在数组越界访问
函数实现代码
split_words函数
char** split_words(const char *text){ if (text == NULL) return NULL; //fun() to count words in given text int words_number = count_words(text); //allocate memory for result array, size: words_number * pointer for each word char **result = (char**)malloc((words_number + 1) * sizeof(char*)); if (result == NULL){ return NULL; } //temporary word variable to save wrods in result array char *word = (char*)malloc(50 * sizeof(char)); if (word == NULL){ destroy(result); return NULL; } int i = 0, next_word=0; while (i < (int)strlen(text)){ while (i < (int)strlen(text)){ if ((*(text + i) >= 'a' && *(text + i) <= 'z') || (*(text + i) >= 'A' && *(text + i) <= 'Z')) break; i++; } int j = 0; while (i < (int)strlen(text)){ if (strchr(" '.,{}[]/\\"-;:"", *(text + i)) != NULL) break; *(word + j) = *(text + i); i++; j++; } if (j > 0){ // add *(word + j) = '\0'; // allocate memory for each word in result array *(result + next_word) = (char*)malloc((strlen(word) + 1) * sizeof(char)); if(*(result + next_word) == NULL){ // destroy(result); free(word); return NULL; } strcpy(*(result + next_word), word); next_word++; } } // add NULL pointer at the end of words array *(result + next_word) = NULL; free(word); return result; }
destroy函数
void destroy(char **words){ if (words == NULL) return; int i = 0; while (*(words + i) != NULL){ free(*(words + i)); i++; } free(words); }
count_words函数
int count_words(const char *text){ if (text == NULL) return 0; int counter = 0, state = 1; for (int i=0 ; *(text + i) != '\0'; i++){ if (strchr(" '.,{}[]/\\"-;:"", *(text + i)) != NULL){ state = 1; } else if (state == 1){ state = 0; counter++; } } return counter; }
问题排查方向
单词判定逻辑不一致
count_words与split_words的分隔符判定逻辑存在隐性差异:split_words的第一个循环仅跳过非字母字符,而count_words只要字符不在分隔符列表就判定为单词起始。若输入包含数字、下划线等非字母但非分隔符的字符,会导致count_words统计的单词数与实际分割出的数量不匹配,进而引发result数组的内存分配大小错误,最终导致越界访问。分隔符字符串转义错误
代码中的分隔符字符串包含\"这类HTML转义字符,而非C语言的正确转义格式(\"),这会导致strchr的匹配逻辑完全错误,使得count_words统计的单词数严重偏离实际值,直接引发result数组的内存分配大小与实际需求不匹配。临时缓冲区长度限制
临时word缓冲区固定分配50字节,若遇到长度超过49的单词,会出现缓冲区写越界,破坏内存中的next_word或result指针值,进而引发后续的越界访问或返回错误指针。内存分配失败时的资源泄漏
当某个单词的malloc失败时,仅释放了临时word缓冲区,未释放已分配的单词指针和result数组本身,不仅会造成内存泄漏,还可能破坏测试环境的内存状态,引发异常报错。重复调用strlen的潜在风险
每次循环都调用strlen(text),虽不会直接引发错误,但会降低效率;若text在运行中被意外修改(尽管参数是const),可能导致循环条件失效。
修复建议
统一单词判定逻辑
提取辅助函数is_delimiter(char c),让count_words和split_words共用同一套分隔符判定逻辑,避免统计与分割的数量差异:static int is_delimiter(char c) { return strchr(" '.,{}[]/\\\"-;:\"", c) != NULL; }修复分隔符字符串转义
将分隔符字符串修正为C语言正确格式:" '.,{}[]/\\\"-;:\"",确保strchr能正确匹配分隔符。动态处理单词缓冲区
移除固定大小的临时缓冲区,改为先定位单词的起始和结束位置,直接分配对应大小的内存存储单词,避免缓冲区越界:// 替换原临时word相关逻辑 while (i < text_len) { // 跳过分隔符 while (i < text_len && is_delimiter(text[i])) i++; if (i >= text_len) break; // 定位单词结束位置 int start = i; while (i < text_len && !is_delimiter(text[i])) i++; int word_len = i - start; // 分配单词内存 result[next_word] = malloc(word_len + 1); if (!result[next_word]) { // 清理已分配资源 for (int k=0; k<next_word; k++) free(result[k]); free(result); return NULL; } strncpy(result[next_word], text + start, word_len); result[next_word][word_len] = '\0'; next_word++; }完善内存分配失败的清理逻辑
当单词malloc失败时,先释放已分配的所有单词指针,再释放result数组,避免内存泄漏和环境破坏。提前计算字符串长度
提前调用一次strlen(text)并保存结果,避免循环中重复调用,提升效率并避免潜在风险。
内容的提问来源于stack exchange,提问作者anocyney_

