You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义C语言ft_split函数拆分多字符串时首元素乱码求助

问题分析与修复

你的ft_split函数出现乱码的核心原因是指针数组的内存分配错误,以及用于统计单词数量的strcnt函数逻辑完全错误,导致内存越界写入,进而出现乱码。

具体错误点

1. strcnt函数逻辑完全错误

你标注该函数的作用是统计非分隔符字符数,但实际逻辑完全偏离:

  • 循环中始终判断*s(即字符串的第一个字符)是否为非分隔符,只要第一个字符不是分隔符,i就会一直递增到整个字符串的长度,返回的是整个字符串的总长度,而非你需要的单词数量。
  • 即使你原本想统计非分隔符字符数,逻辑也错误,应该判断s[i]而非固定的*s。

2. 指针数组的内存分配错误

在ft_split中,你为存储拆分后单词的指针数组strs分配内存时:

strs = malloc (strcnt (s, c) + 1);

这里的问题有两个:

  • 你需要的是存储单词数+1个char*指针的空间,而不是字节数,必须乘以sizeof(char*)才能得到正确的内存大小。
  • strcnt返回的是错误的数值(字符串总长度),导致分配的空间要么过大要么过小,当空间不足时,写入指针会越界覆盖其他内存,引发乱码。

修复后的完整代码

#include <stdio.h>
#include <stdlib.h>

int is_sep(const char s, char c) //Return 1 if the current letter is a separator
{
    return (s == c);
}

size_t wrd_len(const char *s, char c) //Return the length of the word
{
    size_t i = 0;
    while (s[i] != '\0' && !is_sep(s[i], c))
        i++;
    return (i);
}

size_t word_count(const char *s, char c) //统计单词数量
{
    size_t count = 0;
    int in_word = 0;
    while (*s)
    {
        if (is_sep(*s, c))
        {
            in_word = 0;
        }
        else if (!in_word)
        {
            in_word = 1;
            count++;
        }
        s++;
    }
    return count;
}

char *word(const char *s, char c) //Return the word without the separator
{
    size_t wdlen = wrd_len(s, c);
    char *wd = malloc(wdlen + 1);
    if (!wd)
        return (NULL);
    for (size_t i = 0; i < wdlen; i++)
    {
        wd[i] = s[i];
    }
    wd[wdlen] = '\0';
    return (wd);
}

char **ft_split(const char *s, char c)
{
    size_t cnt = word_count(s, c);
    char **strs = malloc((cnt + 1) * sizeof(char*));
    if (!strs)
        return (NULL);
    size_t i = 0;
    while (*s != '\0')
    {
        while (is_sep(*s, c) && *s)
            s++;
        if (*s)
        {
            strs[i] = word(s, c);
            i++;
        }
        while (!is_sep(*s, c) && *s)
            s++;
    }
    strs[i] = NULL;
    return (strs);
}

int main (void)
{
    int i = 0;
    const char test[] = "How are you ? I'm fine !";
    char sep = ' ';
    char **split = ft_split(test, sep);
    while (split[i])
    {
        printf("split %d : %s\n",i,split[i]);
        free(split[i]); // 别忘了释放每个单词的内存
        i++;
    }
    free(split); // 释放指针数组的内存
}

修复说明

  1. 重写了word_count函数(原strcnt),正确统计字符串中的单词数量:通过标记是否处于单词中,遇到非分隔符且不在单词内时,单词计数加1。
  2. 修正了ft_split中指针数组的内存分配:使用(单词数+1)*sizeof(char*)来分配足够的空间存储所有单词指针和末尾的NULL。
  3. 优化了部分代码的写法(比如is_sep直接返回表达式结果,使用size_t替代int更符合标准库的类型规范)。
  4. 添加了内存释放的代码,避免内存泄漏。

测试修复后的代码,输入原测试用例将得到正确的输出:

split 0 : How
split 1 : are
split 2 : you
split 3 : ?
split 4 : I'm
split 5 : fine
split 6 : !

内容的提问来源于stack exchange,提问作者nokosse

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 10:30:58