使用结构体读取带引号CSV文件时触发zsh: trace trap错误排查
带引号CSV解析的C语言
zsh: trace trap错误排查 问题场景
需要解析带引号的CSV文件,示例行如下:
"2015","Canada (excluding territories)","","35 to 49 years","Males","Diabetes","Percent","Percent","239","units","0","v110789429","1.4.2.10.4","4.9","","","","1"
但运行下方C语言代码时,持续返回错误zsh: trace trap,需排查原因并修复。
错误原因分析
1. CSV解析逻辑完全错误
代码中while (strtok(buf, ","))的用法完全不符合strtok的使用规范:strtok的作用是分割字符串并返回子串,但这里既没有接收返回值,也没有用它来定位字段,反而直接通过charPointer遍历buf,逻辑混乱导致内存越界访问——这是zsh: trace trap(内存访问错误)的核心诱因。
2. 行读取方式错误
使用fscanf("%s", buf)读取CSV行时,%s会在空格或换行处停止读取,但CSV字段(比如"Canada (excluding territories)")包含空格,导致buf只能读到字段的前半部分,后续处理全部错乱。
3. 无边界检查的数组操作
- 写入
item、rawItem时,未检查itemPointer、rawItemPointer是否超过数组长度,极易触发缓冲区溢出。 - 结构体中的
geo[25]等数组,当字段内容长度超过容量时,strcpy会直接溢出缓冲区,引发内存错误。
4. 文件读取流程混乱
开头的fscanf(file, "%s", buf)会吃掉第一行的部分内容,随后while (fgetc(file) != EOF)又跳过一个字符,导致行读取错位,大量内容被遗漏或截断。
修复后的代码
以下是重新实现的CSV解析逻辑,针对带引号的字段做了正确处理,同时添加了边界检查:
#include <stdio.h> #include <stdlib.h> #include <string.h> #define MAX_RECORDS 211 #define MAX_LINE_LEN 4096 #define MAX_FIELD_LEN 100 typedef struct question1 { int ref_date; char geo[26]; // 留1位给字符串终止符 char age_group[51]; char sex[11]; double value; } question1; // 从当前位置读取一个带引号的CSV字段 int read_quoted_field(char *buf, int *pos, char *out, int out_len) { int out_pos = 0; // 跳过开头的引号 if (buf[*pos] != '"') return 0; (*pos)++; // 读取到下一个引号(兼容空字段"") while (buf[*pos] != '\0' && buf[*pos] != '"') { if (out_pos >= out_len - 1) break; // 防止缓冲区溢出 out[out_pos++] = buf[*pos]; (*pos)++; } out[out_pos] = '\0'; // 跳过结尾的引号 if (buf[*pos] == '"') (*pos)++; // 跳过字段后的逗号(如果存在) if (buf[*pos] == ',') (*pos)++; return 1; } int main() { FILE *file = fopen("dummy.csv", "r"); if (file == NULL) { printf("文件打开失败\n"); return 1; } question1 arr[MAX_RECORDS]; int records = 0; char buf[MAX_LINE_LEN]; // 跳过表头(如果CSV无表头可删除此行) fgets(buf, MAX_LINE_LEN, file); while (fgets(buf, MAX_LINE_LEN, file) != NULL && records < MAX_RECORDS) { // 去掉行尾的换行符 buf[strcspn(buf, "\n")] = '\0'; question1 temp = {0}; // 初始化结构体成员 int pos = 0; int col = 0; // 读取第0列:ref_date if (read_quoted_field(buf, &pos, buf, MAX_FIELD_LEN)) { sscanf(buf, "%d", &temp.ref_date); } col++; // 读取第1列:geo read_quoted_field(buf, &pos, temp.geo, sizeof(temp.geo)); col++; // 跳过第2列(空字段) read_quoted_field(buf, &pos, buf, MAX_FIELD_LEN); col++; // 读取第3列:age_group read_quoted_field(buf, &pos, temp.age_group, sizeof(temp.age_group)); col++; // 读取第4列:sex read_quoted_field(buf, &pos, temp.sex, sizeof(temp.sex)); col++; // 跳过第5到11列的无关字段 while (col < 12) { read_quoted_field(buf, &pos, buf, MAX_FIELD_LEN); col++; } // 读取第12列:value if (read_quoted_field(buf, &pos, buf, MAX_FIELD_LEN)) { sscanf(buf, "%lf", &temp.value); } arr[records++] = temp; // 可选:打印验证读取结果 printf("日期:%d,地区:%s,年龄组:%s,性别:%s,数值:%.1lf\n", temp.ref_date, temp.geo, temp.age_group, temp.sex, temp.value); } fclose(file); return 0; }
修复说明
- 用
fgets逐行读取CSV,确保完整读取包含空格的行内容。 - 实现
read_quoted_field函数专门处理带引号的CSV字段,正确跳过引号和分隔逗号。 - 所有数组操作添加边界检查,彻底避免缓冲区溢出。
- 结构体数组长度预留字符串终止符位置,避免
strcpy溢出风险。 - 明确跳过不需要的列,逻辑清晰易维护。
内容的提问来源于stack exchange,提问作者MamiCodes
相关产品推荐
相关产品推荐

