You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++20 WSL2 Ubuntu环境下UTF-8 Unicode字符串读取与长度计算问询

C++20 环境下UTF-8文本处理方案

1 计算UTF-8字符串的Unicode字符总数

注意:std::string::size()返回的是UTF-8编码的字节总长度,不是Unicode字符个数,因为UTF-8是变长编码,单个Unicode字符占1~4字节不等。
你可以直接通过遍历UTF-8字节流统计码点数量,不需要依赖第三方库或者编码转换接口,性能更高,实现逻辑如下:

  • 按UTF-8编码规则判断每个字符的字节长度,逐个跳过当前字符的所有字节
  • 遍历完成后统计的总次数就是Unicode字符总个数

2 读取UTF-8文件并正常打印输出

WSL2 Ubuntu的默认终端为UTF-8编码,只要输出标准UTF-8字节流即可正常渲染字符,不需要额外转码。读取文件时以二进制模式打开即可,避免系统默认编码转换导致乱码。

完整可运行代码示例

#include <fstream>
#include <string>
#include <iostream>

// 统计UTF-8字符串的Unicode字符总数
size_t count_utf8_codepoints(const std::string& utf8_str) {
    size_t count = 0;
    size_t i = 0;
    const size_t len = utf8_str.size();
    while (i < len) {
        unsigned char c = static_cast<unsigned char>(utf8_str[i]);
        if (c < 0x80) i += 1;        // 单字节字符
        else if (c < 0xE0) i += 2;   // 双字节字符
        else if (c < 0xF0) i += 3;   // 三字节字符
        else i += 4;                 // 四字节字符
        count++;
    }
    return count;
}

int main() {
    // 二进制模式读取UTF-8文件,避免系统默认编码转换
    std::ifstream fin("test.txt", std::ios::binary);
    std::string utf8_content((std::istreambuf_iterator<char>(fin)), std::istreambuf_iterator<char>());

    // 输出Unicode字符总数
    std::cout << "Unicode字符总个数:" << count_utf8_codepoints(utf8_content) << std::endl;

    // 直接输出UTF-8字节流到终端,WSL2默认支持UTF-8渲染
    std::cout << "文件内容打印:\n" << utf8_content << std::endl;

    return 0;
}

编译运行命令

# 编译,指定C++20标准
g++ -std=c++20 utf8_test.cpp -o utf8_test
# 运行
./utf8_test

内容的提问来源于stack exchange,提问作者spaL

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 20:15:00