You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CSV(逗号分隔)文件加载数据函数异常排查求助

问题分析:CSV文件读取异常的原因排查

我在Visual Studio中为EEG睡眠研究实验室编写C语言代码,需加载一个包含50行3000列的CSV(逗号分隔)Excel文件,每行对应一条含3000个数据点的时间序列信号。但运行代码后,统计得到的行数为500、列数为316,读取的数据值也完全错误。此前一学期使用相同的load_data、num_rows_in_file、num_cols_in_file函数均无异常,请问代码问题出在哪里?

原代码

int main(int argc, char* argv[]) {

    FILE* file = fopen("EEG_SleepData_30sec_100Hz.csv", "r");
    if (file == NULL) {
        perror("Error opening file");
        return EXIT_FAILURE;
    }

    //Loads data in from file name specified
    int num_signals = num_rows_in_file(file);
    int signal_length = num_cols_in_file(file);
    printf("number of rows = %d  number of columns = %d\n", num_signals, signal_length);

    double** dataset = load_data(file, num_signals, signal_length);
    // Print the entire dataset
    for (int i = 0; i < num_signals; i++) {
        for (int j = 0; j < signal_length; j++) {
            printf("%lf ", dataset[i][j]);
        }
        printf("\n");
    }
    
    return 0;
}


double** load_data(FILE* file, int numrows, int numcols) {

    if (file != NULL) {
        double** dataset = (double**)calloc(numrows, sizeof(double*)); // Allocate each of our row pointers.

        if (dataset == NULL) {
            return NULL;
        }

        for (int i = 0; i < numrows; i++) {
            dataset[i] = (double*)calloc(numcols, sizeof(double)); // Allocate our columns.
            if (dataset[i] == NULL) {
                return NULL;
            }
        }

        for (int i = 0; i < numrows; i++) {
            for (int j = 0; j < numcols; j++) {
                fscanf(file, "%lf,", &dataset[i][j]);
            }
        }
        return dataset;
    }
    else {
        fprintf(stderr, "Unable to find file! Ensure it is in the Debug directory.");
        return NULL;
    }

}


int num_cols_in_file(FILE* file) {
    int numcols = 0;
    if (file) {
        char buf[3000]; // Make a buffer we'll use to grab a whole row.

        if (fgets(buf, sizeof(buf), file) != NULL) {
            // Tokenize our buffer, looking for how many columns we have (aka how many tokens we can create)
            char* token;
            char* next_token = NULL;

            token = strtok(buf, ", \n\r\t", &next_token);  // Include commas and additional whitespace as delimiters

            while (token != NULL) {
                token = strtok(NULL, ", \n\r\t", &next_token);
                numcols++;
            }

            rewind(file); // Reset our position to the beginning of the file.

            return numcols;
        }
        else {
            fprintf(stderr, "Failed to read first row.\n");
            return 0;
        }
    }
    else {
        fprintf(stderr, "File is unopened. Numcols only works on opened files.");
        return 0;
    }
}


int num_rows_in_file(FILE* file) {
    int numrows = 0;
    if (file) {
        char buf[3000]; // Make a buffer we'll use to grab a whole row.

        while (fgets(buf, sizeof(buf), file) != NULL) {
            numrows++;

        }

        rewind(file); // Reset our position to the beginning of the file.
        return numrows;
    }
    else {
        fprintf(stderr, "File is unopened. Numrows only works on opened files.");
        return 0;
    }
}

问题根源

  1. 缓冲区长度严重不足
    num_cols_in_file和num_rows_in_file中使用的buf[3000]仅能容纳3000字节,但每行有3000个数据点,单个数据点(含逗号分隔符)至少需要5-10字节,整行长度远超过3000字节。fgets会截断行内容,导致:

    • num_rows_in_file把截断的剩余内容当成新行,行数统计被放大10倍(50→500);
    • num_cols_in_file只能统计到截断部分的列数,出现316的错误结果。
  2. 列数统计逻辑错误
    num_cols_in_file中,第一次调用strtok拿到第一个列的token后,循环里先调用strtok获取下一个token再执行numcols++,导致第一个列未被计数,实际列数会少1。

  3. fscanf读取的格式兼容性问题
    load_data中使用%lf,格式串,要求每个数据后必须跟逗号,但CSV文件的行尾通常没有逗号。这会导致每行最后一个数据读取失败,后续读取会错位,进一步加剧数据混乱。

修正方案

  1. 扩大缓冲区或动态分配
    将buf的大小改为足够容纳整行的长度,比如char buf[65536];,或者根据文件行动态分配内存。

  2. 修正列数统计逻辑
    调整num_cols_in_file的计数顺序:

    token = strtok(buf, ", \n\r\t", &next_token);
    while (token != NULL) {
        numcols++;
        token = strtok(NULL, ", \n\r\t", &next_token);
    }
    
  3. 改用逐行解析的方式读取数据
    在load_data中,先读取整行,再用strtok拆分每个数据点,避免fscanf的格式问题:

    char line_buf[65536];
    for (int i = 0; i < numrows; i++) {
        if (fgets(line_buf, sizeof(line_buf), file) == NULL) break;
        char* token = strtok(line_buf, ", \n\r\t");
        int j = 0;
        while (token != NULL && j < numcols) {
            dataset[i][j++] = atof(token);
            token = strtok(NULL, ", \n\r\t");
        }
    }
    

内容的提问来源于stack exchange,提问作者MacKenna Bochnak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 15:35:33