C语言读取含逗号的CSV文件异常求助:游戏数据解析错位
解决带引号逗号的CSV字段分割问题
strtok的核心问题是它仅按单一分隔符(逗号)分割字符串,完全忽略引号的上下文,导致字段内的逗号被误判为列分隔符。要解决这个问题,你需要实现一个支持识别引号包裹字段的CSV解析逻辑,而不是依赖strtok。
自定义CSV字段解析函数
下面是一个轻量的解析函数,能正确处理带双引号包裹的字段,同时兼容普通字段:
#include <stdio.h> #include <stdlib.h> #include <string.h> #include <ctype.h> // 解析单个CSV字段,返回字段内容的副本(需手动free),同时移动指针到下一个字段的起始位置 char* parse_csv_field(char** ptr) { char* start = *ptr; char* field; int len; // 跳过字段开头的空白字符(可选,根据你的CSV格式调整) while (isspace((unsigned char)*start)) start++; if (*start == '"') { // 处理带引号的字段:找到下一个非转义的双引号 start++; // 跳过开头的引号 char* end = start; while (*end != '\0' && !(*end == '"' && *(end+1) != '"')) { // 处理转义引号(如果CSV用""表示单个") if (*end == '"' && *(end+1) == '"') { memmove(end, end+1, strlen(end)); // 移除重复的引号 continue; } end++; } len = end - start; field = malloc(len + 1); strncpy(field, start, len); field[len] = '\0'; // 移动指针到引号后的逗号/行尾 *ptr = (*end == '"') ? end + 1 : end; } else { // 处理普通字段:找到下一个逗号或行尾 char* end = strchr(start, ','); if (!end) end = start + strlen(start); len = end - start; field = malloc(len + 1); strncpy(field, start, len); field[len] = '\0'; *ptr = end; } // 跳过字段后的逗号,准备下一个字段 if (**ptr == ',') (*ptr)++; return field; }
完整示例代码(结合你的需求)
以下是整合了解析函数、数据读取、排序和Top10输出的完整代码:
typedef struct { char* title; int score; char* genre; int year; } Game; // 排序用的比较函数:按分数降序排列 int compare_games(const void* a, const void* b) { const Game* gameA = (const Game*)a; const Game* gameB = (const Game*)b; return gameB->score - gameA->score; } int main() { FILE* fp = fopen("games.csv", "r"); if (!fp) { perror("Failed to open file"); return 1; } char line[1024]; Game* games = NULL; int game_count = 0; // 跳过CSV表头 fgets(line, sizeof(line), fp); // 逐行读取并解析数据 while (fgets(line, sizeof(line), fp)) { // 移除换行符 line[strcspn(line, "\n")] = '\0'; char* ptr = line; Game new_game; new_game.title = parse_csv_field(&ptr); new_game.score = atoi(parse_csv_field(&ptr)); new_game.genre = parse_csv_field(&ptr); new_game.year = atoi(parse_csv_field(&ptr)); // 动态扩容数组 games = realloc(games, (game_count + 1) * sizeof(Game)); games[game_count++] = new_game; } fclose(fp); // 按分数排序 qsort(games, game_count, sizeof(Game), compare_games); // 输出Top10 printf("Top 10 Highest Rated Games:\n"); int top_count = (game_count < 10) ? game_count : 10; for (int i = 0; i < top_count; i++) { printf("%d. %s - Score: %d, Genre: %s, Year: %d\n", i+1, games[i].title, games[i].score, games[i].genre, games[i].year); } // 释放内存 for (int i = 0; i < game_count; i++) { free(games[i].title); free(games[i].genre); } free(games); return 0; }
关键说明
- 解析逻辑:函数会先判断字段是否以双引号开头,如果是则将下一个双引号作为字段结束(同时处理
""转义为单个"的情况);否则以逗号作为分隔符。 - 内存管理:解析函数返回的字段是动态分配的,使用完后必须手动释放,避免内存泄漏。
- 兼容性:这个逻辑兼容普通无引号字段,也能处理带内部逗号的引号字段,符合标准CSV格式的基本规则。
内容的提问来源于stack exchange,提问作者Sukhman
相关产品推荐
相关产品推荐

