You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C语言中fwscanf无法正确读取UTF-8格式CSV文件的问题

使用C标准库的fwscanf读取UTF-8 CSV文件的问题解决

本程序仅可使用C标准库。

我尝试在C语言中使用fwscanf读取UTF-8编码的CSV文件,但读取过程中遇到问题。该文件每行包含一个字符串和一个浮点数值,以逗号分隔。以下是复现问题的最小示例:

#include <stdio.h>
#include <wchar.h>
#include <locale.h>

#define MAX_STRING_LENGTH 31

int main() {
    setlocale(LC_ALL, "en_US.UTF-8");
    FILE *file = fopen("input.csv", "r, ccs=UTF-8");
    if (file == NULL) {
        fwprintf(stderr, L"Error opening file.\n");
        return 1;
    }

    wchar_t string[MAX_STRING_LENGTH];
    float frequency;
    int row = 0;

    while (!feof(file)) {
        row++;
        int result = fwscanf(file, L"%30[^,],%f,", string, &frequency);
        
        if (result == 2) {
            wprintf(L"Row %d: String = '%ls', Frequency = %.4f\n", row, string, frequency);
        } else if (result == 1) {
            wprintf(L"Row %d: String = '%ls', Frequency not read\n", row, string);
        } else if (result == EOF) {
            break;
        } else {
            wprintf(L"Error reading row %d\n", row);
            wchar_t c;
            // Skip the rest of the line
            while ((c = fgetwc(file)) != L'\n' && c != WEOF);
        }
    }

    fclose(file);
    return 0;
}

示例input.csv内容:

hello,1.0000
world,0.5000
how,0.7500
are,0.2500
you,1.0000
?,0.5000

预期输出:

Row 1: String = 'hello', Frequency = 1.0000
Row 2: String = 'world', Frequency = 0.5000
Row 3: String = 'how', Frequency = 0.7500
Row 4: String = 'are', Frequency = 0.2500
Row 5: String = 'you', Frequency = 1.0000
Row 6: String = '?', Frequency = 0.5000

遇到的问题:fwscanf无法正确读取文件,要么读取到错误值,要么完全读取失败。尝试过调整区域设置和文件打开模式,但问题仍未解决。


问题分析与修正

1. 格式字符串错误

原格式串L"%30[^,],%f,"末尾多了一个逗号,而CSV每行的结构是字符串,数值\n,没有末尾逗号。第一次读取后,文件指针会停在数值后的换行符位置,下一次读取时%30[^,]会读取换行符,导致匹配失败。需移除末尾的逗号,改为L"%30[^,],%f"。

2. feof循环陷阱

while (!feof(file))的逻辑错误:feof仅在读取失败后才会置位,因此读取完最后一行后仍会进入循环执行一次,导致错误的行计数和读取行为。应直接用fwscanf的返回值控制循环。

3. 区域设置未做检查

setlocale可能调用失败(比如系统未安装对应locale),此时宽字符函数无法正确处理UTF-8,需添加成功检查。

4. 换行符残留处理

读取完每行的字符串和数值后,需手动跳过剩余的换行符,避免残留字符影响下一次读取。

修正后的代码

#include <stdio.h>
#include <wchar.h>
#include <locale.h>

#define MAX_STRING_LENGTH 31

int main() {
    // 检查区域设置是否成功
    if (!setlocale(LC_ALL, "en_US.UTF-8")) {
        fwprintf(stderr, L"Failed to set locale.\n");
        return 1;
    }

    FILE *file = fopen("input.csv", "r, ccs=UTF-8");
    if (file == NULL) {
        fwprintf(stderr, L"Error opening file.\n");
        return 1;
    }

    wchar_t string[MAX_STRING_LENGTH];
    float frequency;
    int row = 0;
    int result;

    // 用fwscanf返回值控制循环,避免feof陷阱
    while ((result = fwscanf(file, L"%30[^,],%f", string, &frequency)) != EOF) {
        row++;
        if (result == 2) {
            wprintf(L"Row %d: String = '%ls', Frequency = %.4f\n", row, string, frequency);
        } else if (result == 1) {
            wprintf(L"Row %d: String = '%ls', Frequency not read\n", row, string);
        } else {
            wprintf(L"Error reading row %d\n", row);
        }
        // 跳过当前行剩余字符(含换行)
        wchar_t c;
        while ((c = fgetwc(file)) != L'\n' && c != WEOF);
    }

    fclose(file);
    return 0;
}

额外说明

  • 修正后的代码会正确解析UTF-8字符串(比如示例中的?),因为已成功设置UTF-8 locale,fwscanf会自动将文件中的UTF-8字节序列转换为宽字符。
  • 如果CSV字符串中包含逗号,%[^,]会提前终止读取,这种场景需要更复杂的CSV解析逻辑,但符合当前示例的需求。

内容的提问来源于stack exchange,提问作者iPc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 16:30:09