自定义C语言ft_split函数拆分多字符串时首元素乱码求助
问题分析与修复
你的ft_split函数出现乱码的核心原因是指针数组的内存分配错误,以及用于统计单词数量的strcnt函数逻辑完全错误,导致内存越界写入,进而出现乱码。
具体错误点
1. strcnt函数逻辑完全错误
你标注该函数的作用是统计非分隔符字符数,但实际逻辑完全偏离:
- 循环中始终判断
*s(即字符串的第一个字符)是否为非分隔符,只要第一个字符不是分隔符,i就会一直递增到整个字符串的长度,返回的是整个字符串的总长度,而非你需要的单词数量。 - 即使你原本想统计非分隔符字符数,逻辑也错误,应该判断
s[i]而非固定的*s。
2. 指针数组的内存分配错误
在ft_split中,你为存储拆分后单词的指针数组strs分配内存时:
strs = malloc (strcnt (s, c) + 1);
这里的问题有两个:
- 你需要的是存储
单词数+1个char*指针的空间,而不是字节数,必须乘以sizeof(char*)才能得到正确的内存大小。 strcnt返回的是错误的数值(字符串总长度),导致分配的空间要么过大要么过小,当空间不足时,写入指针会越界覆盖其他内存,引发乱码。
修复后的完整代码
#include <stdio.h> #include <stdlib.h> int is_sep(const char s, char c) //Return 1 if the current letter is a separator { return (s == c); } size_t wrd_len(const char *s, char c) //Return the length of the word { size_t i = 0; while (s[i] != '\0' && !is_sep(s[i], c)) i++; return (i); } size_t word_count(const char *s, char c) //统计单词数量 { size_t count = 0; int in_word = 0; while (*s) { if (is_sep(*s, c)) { in_word = 0; } else if (!in_word) { in_word = 1; count++; } s++; } return count; } char *word(const char *s, char c) //Return the word without the separator { size_t wdlen = wrd_len(s, c); char *wd = malloc(wdlen + 1); if (!wd) return (NULL); for (size_t i = 0; i < wdlen; i++) { wd[i] = s[i]; } wd[wdlen] = '\0'; return (wd); } char **ft_split(const char *s, char c) { size_t cnt = word_count(s, c); char **strs = malloc((cnt + 1) * sizeof(char*)); if (!strs) return (NULL); size_t i = 0; while (*s != '\0') { while (is_sep(*s, c) && *s) s++; if (*s) { strs[i] = word(s, c); i++; } while (!is_sep(*s, c) && *s) s++; } strs[i] = NULL; return (strs); } int main (void) { int i = 0; const char test[] = "How are you ? I'm fine !"; char sep = ' '; char **split = ft_split(test, sep); while (split[i]) { printf("split %d : %s\n",i,split[i]); free(split[i]); // 别忘了释放每个单词的内存 i++; } free(split); // 释放指针数组的内存 }
修复说明
- 重写了
word_count函数(原strcnt),正确统计字符串中的单词数量:通过标记是否处于单词中,遇到非分隔符且不在单词内时,单词计数加1。 - 修正了
ft_split中指针数组的内存分配:使用(单词数+1)*sizeof(char*)来分配足够的空间存储所有单词指针和末尾的NULL。 - 优化了部分代码的写法(比如
is_sep直接返回表达式结果,使用size_t替代int更符合标准库的类型规范)。 - 添加了内存释放的代码,避免内存泄漏。
测试修复后的代码,输入原测试用例将得到正确的输出:
split 0 : How split 1 : are split 2 : you split 3 : ? split 4 : I'm split 5 : fine split 6 : !
内容的提问来源于stack exchange,提问作者nokosse
相关产品推荐
相关产品推荐

