CSV(逗号分隔)文件加载数据函数异常排查求助
我在Visual Studio中为EEG睡眠研究实验室编写C语言代码,需加载一个包含50行3000列的CSV(逗号分隔)Excel文件,每行对应一条含3000个数据点的时间序列信号。但运行代码后,统计得到的行数为500、列数为316,读取的数据值也完全错误。此前一学期使用相同的load_data、num_rows_in_file、num_cols_in_file函数均无异常,请问代码问题出在哪里?
原代码
int main(int argc, char* argv[]) { FILE* file = fopen("EEG_SleepData_30sec_100Hz.csv", "r"); if (file == NULL) { perror("Error opening file"); return EXIT_FAILURE; } //Loads data in from file name specified int num_signals = num_rows_in_file(file); int signal_length = num_cols_in_file(file); printf("number of rows = %d number of columns = %d\n", num_signals, signal_length); double** dataset = load_data(file, num_signals, signal_length); // Print the entire dataset for (int i = 0; i < num_signals; i++) { for (int j = 0; j < signal_length; j++) { printf("%lf ", dataset[i][j]); } printf("\n"); } return 0; } double** load_data(FILE* file, int numrows, int numcols) { if (file != NULL) { double** dataset = (double**)calloc(numrows, sizeof(double*)); // Allocate each of our row pointers. if (dataset == NULL) { return NULL; } for (int i = 0; i < numrows; i++) { dataset[i] = (double*)calloc(numcols, sizeof(double)); // Allocate our columns. if (dataset[i] == NULL) { return NULL; } } for (int i = 0; i < numrows; i++) { for (int j = 0; j < numcols; j++) { fscanf(file, "%lf,", &dataset[i][j]); } } return dataset; } else { fprintf(stderr, "Unable to find file! Ensure it is in the Debug directory."); return NULL; } } int num_cols_in_file(FILE* file) { int numcols = 0; if (file) { char buf[3000]; // Make a buffer we'll use to grab a whole row. if (fgets(buf, sizeof(buf), file) != NULL) { // Tokenize our buffer, looking for how many columns we have (aka how many tokens we can create) char* token; char* next_token = NULL; token = strtok(buf, ", \n\r\t", &next_token); // Include commas and additional whitespace as delimiters while (token != NULL) { token = strtok(NULL, ", \n\r\t", &next_token); numcols++; } rewind(file); // Reset our position to the beginning of the file. return numcols; } else { fprintf(stderr, "Failed to read first row.\n"); return 0; } } else { fprintf(stderr, "File is unopened. Numcols only works on opened files."); return 0; } } int num_rows_in_file(FILE* file) { int numrows = 0; if (file) { char buf[3000]; // Make a buffer we'll use to grab a whole row. while (fgets(buf, sizeof(buf), file) != NULL) { numrows++; } rewind(file); // Reset our position to the beginning of the file. return numrows; } else { fprintf(stderr, "File is unopened. Numrows only works on opened files."); return 0; } }
问题根源
缓冲区长度严重不足
num_cols_in_file和num_rows_in_file中使用的buf[3000]仅能容纳3000字节,但每行有3000个数据点,单个数据点(含逗号分隔符)至少需要5-10字节,整行长度远超过3000字节。fgets会截断行内容,导致:num_rows_in_file把截断的剩余内容当成新行,行数统计被放大10倍(50→500);num_cols_in_file只能统计到截断部分的列数,出现316的错误结果。
列数统计逻辑错误
num_cols_in_file中,第一次调用strtok拿到第一个列的token后,循环里先调用strtok获取下一个token再执行numcols++,导致第一个列未被计数,实际列数会少1。fscanf读取的格式兼容性问题load_data中使用%lf,格式串,要求每个数据后必须跟逗号,但CSV文件的行尾通常没有逗号。这会导致每行最后一个数据读取失败,后续读取会错位,进一步加剧数据混乱。
修正方案
扩大缓冲区或动态分配
将buf的大小改为足够容纳整行的长度,比如char buf[65536];,或者根据文件行动态分配内存。修正列数统计逻辑
调整num_cols_in_file的计数顺序:token = strtok(buf, ", \n\r\t", &next_token); while (token != NULL) { numcols++; token = strtok(NULL, ", \n\r\t", &next_token); }改用逐行解析的方式读取数据
在load_data中,先读取整行,再用strtok拆分每个数据点,避免fscanf的格式问题:char line_buf[65536]; for (int i = 0; i < numrows; i++) { if (fgets(line_buf, sizeof(line_buf), file) == NULL) break; char* token = strtok(line_buf, ", \n\r\t"); int j = 0; while (token != NULL && j < numcols) { dataset[i][j++] = atof(token); token = strtok(NULL, ", \n\r\t"); } }
内容的提问来源于stack exchange,提问作者MacKenna Bochnak

