使用strtok实现单词生成器遇段错误,如何处理不确定单词数?
C语言单词生成器段错误分析与动态扩容解决方案
段错误的核心原因
你代码里的words数组是用固定大小NUM_WORDS分配的,当文本中的实际单词数量超过这个值时,words[num_words]会越界访问非法内存,直接触发Segmentation Fault。这是处理不确定数量元素时的典型问题——固定容器无法适配可变长度的输入。
解决思路与修正代码
要处理数量不确定的单词,必须用动态扩容的方式管理存储单词的指针数组,同时优化单词处理逻辑避免其他潜在问题:
关键优化点
- 动态扩容数组:初始分配一个合理的小容量,当元素数量达到容量上限时,用
realloc扩容(通常每次翻倍,减少扩容次数)。 - 使用线程安全的
strtok_r替代strtok:strtok依赖全局静态变量,多线程或嵌套调用时会出问题,strtok_r更可靠。 - 复制原文本:
strtok系列函数会修改输入字符串,复制一份避免破坏原始数据。 - 精确复制单词:根据处理后的长度复制内容,避免保留多余空白字符,同时跳过空单词(比如连续空白产生的无效内容)。
修正后的完整代码
#include <stdio.h> #include <stdlib.h> #include <string.h> #include <ctype.h> char** split_words(const char* text, int* out_num_words) { if (text == NULL || out_num_words == NULL) { return NULL; } int capacity = 8; // 初始容量,可按需调整 char** words = malloc(sizeof(char*) * capacity); if (words == NULL) { return NULL; } int num_words = 0; // 复制原文本,防止strtok_r修改原始数据 char* text_copy = strdup(text); if (text_copy == NULL) { free(words); return NULL; } char* save_ptr; // 用所有空白字符作为分隔符,适配空格、制表符、换行等 char* word = strtok_r(text_copy, " \t\n\r\f\v", &save_ptr); while (word != NULL) { // 去掉单词尾部的空白字符 int len = strlen(word); while (len > 0 && isspace((unsigned char)word[len - 1])) { len--; } // 跳过空单词(比如连续空白产生的无效内容) if (len == 0) { word = strtok_r(NULL, " \t\n\r\f\v", &save_ptr); continue; } // 检查是否需要扩容 if (num_words >= capacity) { capacity *= 2; char** temp = realloc(words, sizeof(char*) * capacity); if (temp == NULL) { // 内存分配失败,清理已分配资源 for (int i = 0; i < num_words; i++) { free(words[i]); } free(words); free(text_copy); return NULL; } words = temp; } // 为处理后的单词分配内存并复制 words[num_words] = malloc(len + 1); if (words[num_words] == NULL) { // 清理资源 for (int i = 0; i < num_words; i++) { free(words[i]); } free(words); free(text_copy); return NULL; } strncpy(words[num_words], word, len); words[num_words][len] = '\0'; // 手动添加字符串结束符 num_words++; word = strtok_r(NULL, " \t\n\r\f\v", &save_ptr); } free(text_copy); // 可选:缩容到实际需要的大小,节省内存 char** final_words = realloc(words, sizeof(char*) * num_words); if (final_words != NULL) { words = final_words; } *out_num_words = num_words; return words; } // 使用示例 int main() { const char* text = "some text without comma and extra spaces\twith tab\nnew line word"; int num_words; char** words = split_words(text, &num_words); if (words == NULL) { fprintf(stderr, "Failed to split words\n"); return 1; } for (int i = 0; i < num_words; i++) { printf("Word %d: %s\n", i+1, words[i]); free(words[i]); } free(words); return 0; }
使用说明
- 调用
split_words时传入文本和一个用于接收单词数量的指针,函数返回存储单词的指针数组。 - 使用完毕后要逐个释放每个单词的内存,最后释放数组本身,避免内存泄漏。
内容的提问来源于stack exchange,提问作者kenjsch
相关产品推荐
相关产品推荐

