You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何调用std::isgraph时会在std::use_facet中抛出bad_cast异常?

Why does std::isgraph<char32_t> throw std::bad_cast with en_US.utf8 locale?

The root cause here is that your Ubuntu system's en_US.utf8 locale does not include a std::ctype<char32_t> facet—and std::isgraph<char32_t> depends entirely on this facet to determine if a character is graphical.

Let’s break down what’s happening:

  • When you call std::isgraph(code_point, loc) with a char32_t argument, the function internally uses std::use_facet<std::ctype<char32_t>>(loc) to fetch the character classification facet for char32_t.
  • Glibc (the C standard library used on Ubuntu) only ships with pre-built ctype facets for char and wchar_t by default. Standard locales like en_US.utf8 don’t come with a pre-made ctype<char32_t> facet.
  • As you saw in the use_facet implementation, when the facet index (__i) matches the total number of facets in the locale (_M_facets_size), it means the requested facet doesn’t exist—so it throws std::bad_cast.

You can confirm this with a quick test:

#include <iostream>
#include <locale>

int main() {
    std::locale loc("en_US.utf8");
    std::cout << std::boolalpha << std::has_facet<std::ctype<char32_t>>(loc) << "\n";
    // This prints false on Ubuntu
    return 0;
}

Solutions to fix this:

1. Convert char32_t to wchar_t (simplest for Linux)

On Ubuntu and most Linux systems, wchar_t is a 32-bit type that natively supports UTF-32 (matching char32_t). You can safely cast your code point to wchar_t and use the std::isgraph overload for wchar_t, which works with the existing ctype<wchar_t> facet in en_US.utf8:

#include <locale>
#include <cctype>

bool foo(char32_t code_point) { 
    static std::locale loc("en_US.utf8"); 
    wchar_t wide_code = static_cast<wchar_t>(code_point);
    return std::isgraph(wide_code, loc); 
}

int main() {
    foo('B'); // No exception thrown now
    return 0;
}

2. Create a custom ctype<char32_t> facet

If you need to work directly with char32_t without conversion, you can build a custom ctype<char32_t> facet that delegates to the existing ctype<wchar_t> facet. This reuses the character classification rules from en_US.utf8:

#include <locale>
#include <cctype>
#include <array>

struct ctype_char32 : std::ctype<char32_t> {
    ctype_char32(std::size_t refs = 0) 
        : std::ctype<char32_t>(get_table(), false, refs) {}

private:
    static const mask* get_table() {
        static std::array<mask, table_size> table;
        // Copy classification rules from ctype<wchar_t>
        const auto& wctype = std::use_facet<std::ctype<wchar_t>>(std::locale("en_US.utf8"));
        std::copy(wctype.table(), wctype.table() + std::ctype<wchar_t>::table_size, table.begin());
        return table.data();
    }
};

bool foo(char32_t code_point) { 
    // Attach our custom facet to the en_US.utf8 locale
    static std::locale loc(std::locale("en_US.utf8"), new ctype_char32); 
    return std::isgraph(code_point, loc); 
}

int main() {
    foo('B'); // No exception thrown now
    return 0;
}

Portability note:

This first solution relies on wchar_t being 32-bit (true for Linux, but not Windows—where wchar_t is 16-bit). For cross-platform support, you’d need to handle UTF-16 conversion, but for Ubuntu systems, the first approach is straightforward and reliable.

内容的提问来源于stack exchange,提问作者spraff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:46:40