C语言中fwscanf无法正确读取UTF-8格式CSV文件的问题
使用C标准库的fwscanf读取UTF-8 CSV文件的问题解决
本程序仅可使用C标准库。
我尝试在C语言中使用fwscanf读取UTF-8编码的CSV文件,但读取过程中遇到问题。该文件每行包含一个字符串和一个浮点数值,以逗号分隔。以下是复现问题的最小示例:
#include <stdio.h> #include <wchar.h> #include <locale.h> #define MAX_STRING_LENGTH 31 int main() { setlocale(LC_ALL, "en_US.UTF-8"); FILE *file = fopen("input.csv", "r, ccs=UTF-8"); if (file == NULL) { fwprintf(stderr, L"Error opening file.\n"); return 1; } wchar_t string[MAX_STRING_LENGTH]; float frequency; int row = 0; while (!feof(file)) { row++; int result = fwscanf(file, L"%30[^,],%f,", string, &frequency); if (result == 2) { wprintf(L"Row %d: String = '%ls', Frequency = %.4f\n", row, string, frequency); } else if (result == 1) { wprintf(L"Row %d: String = '%ls', Frequency not read\n", row, string); } else if (result == EOF) { break; } else { wprintf(L"Error reading row %d\n", row); wchar_t c; // Skip the rest of the line while ((c = fgetwc(file)) != L'\n' && c != WEOF); } } fclose(file); return 0; }
示例input.csv内容:
hello,1.0000 world,0.5000 how,0.7500 are,0.2500 you,1.0000 ?,0.5000
预期输出:
Row 1: String = 'hello', Frequency = 1.0000 Row 2: String = 'world', Frequency = 0.5000 Row 3: String = 'how', Frequency = 0.7500 Row 4: String = 'are', Frequency = 0.2500 Row 5: String = 'you', Frequency = 1.0000 Row 6: String = '?', Frequency = 0.5000
遇到的问题:fwscanf无法正确读取文件,要么读取到错误值,要么完全读取失败。尝试过调整区域设置和文件打开模式,但问题仍未解决。
问题分析与修正
1. 格式字符串错误
原格式串L"%30[^,],%f,"末尾多了一个逗号,而CSV每行的结构是字符串,数值\n,没有末尾逗号。第一次读取后,文件指针会停在数值后的换行符位置,下一次读取时%30[^,]会读取换行符,导致匹配失败。需移除末尾的逗号,改为L"%30[^,],%f"。
2. feof循环陷阱
while (!feof(file))的逻辑错误:feof仅在读取失败后才会置位,因此读取完最后一行后仍会进入循环执行一次,导致错误的行计数和读取行为。应直接用fwscanf的返回值控制循环。
3. 区域设置未做检查
setlocale可能调用失败(比如系统未安装对应locale),此时宽字符函数无法正确处理UTF-8,需添加成功检查。
4. 换行符残留处理
读取完每行的字符串和数值后,需手动跳过剩余的换行符,避免残留字符影响下一次读取。
修正后的代码
#include <stdio.h> #include <wchar.h> #include <locale.h> #define MAX_STRING_LENGTH 31 int main() { // 检查区域设置是否成功 if (!setlocale(LC_ALL, "en_US.UTF-8")) { fwprintf(stderr, L"Failed to set locale.\n"); return 1; } FILE *file = fopen("input.csv", "r, ccs=UTF-8"); if (file == NULL) { fwprintf(stderr, L"Error opening file.\n"); return 1; } wchar_t string[MAX_STRING_LENGTH]; float frequency; int row = 0; int result; // 用fwscanf返回值控制循环,避免feof陷阱 while ((result = fwscanf(file, L"%30[^,],%f", string, &frequency)) != EOF) { row++; if (result == 2) { wprintf(L"Row %d: String = '%ls', Frequency = %.4f\n", row, string, frequency); } else if (result == 1) { wprintf(L"Row %d: String = '%ls', Frequency not read\n", row, string); } else { wprintf(L"Error reading row %d\n", row); } // 跳过当前行剩余字符(含换行) wchar_t c; while ((c = fgetwc(file)) != L'\n' && c != WEOF); } fclose(file); return 0; }
额外说明
- 修正后的代码会正确解析UTF-8字符串(比如示例中的
?),因为已成功设置UTF-8 locale,fwscanf会自动将文件中的UTF-8字节序列转换为宽字符。 - 如果CSV字符串中包含逗号,
%[^,]会提前终止读取,这种场景需要更复杂的CSV解析逻辑,但符合当前示例的需求。
内容的提问来源于stack exchange,提问作者iPc
相关产品推荐
相关产品推荐

