为何调用std::isgraph时会在std::use_facet中抛出bad_cast异常?
std::isgraph<char32_t> throw std::bad_cast with en_US.utf8 locale? The root cause here is that your Ubuntu system's en_US.utf8 locale does not include a std::ctype<char32_t> facet—and std::isgraph<char32_t> depends entirely on this facet to determine if a character is graphical.
Let’s break down what’s happening:
- When you call
std::isgraph(code_point, loc)with achar32_targument, the function internally usesstd::use_facet<std::ctype<char32_t>>(loc)to fetch the character classification facet forchar32_t. - Glibc (the C standard library used on Ubuntu) only ships with pre-built
ctypefacets forcharandwchar_tby default. Standard locales likeen_US.utf8don’t come with a pre-madectype<char32_t>facet. - As you saw in the
use_facetimplementation, when the facet index (__i) matches the total number of facets in the locale (_M_facets_size), it means the requested facet doesn’t exist—so it throwsstd::bad_cast.
You can confirm this with a quick test:
#include <iostream> #include <locale> int main() { std::locale loc("en_US.utf8"); std::cout << std::boolalpha << std::has_facet<std::ctype<char32_t>>(loc) << "\n"; // This prints false on Ubuntu return 0; }
Solutions to fix this:
1. Convert char32_t to wchar_t (simplest for Linux)
On Ubuntu and most Linux systems, wchar_t is a 32-bit type that natively supports UTF-32 (matching char32_t). You can safely cast your code point to wchar_t and use the std::isgraph overload for wchar_t, which works with the existing ctype<wchar_t> facet in en_US.utf8:
#include <locale> #include <cctype> bool foo(char32_t code_point) { static std::locale loc("en_US.utf8"); wchar_t wide_code = static_cast<wchar_t>(code_point); return std::isgraph(wide_code, loc); } int main() { foo('B'); // No exception thrown now return 0; }
2. Create a custom ctype<char32_t> facet
If you need to work directly with char32_t without conversion, you can build a custom ctype<char32_t> facet that delegates to the existing ctype<wchar_t> facet. This reuses the character classification rules from en_US.utf8:
#include <locale> #include <cctype> #include <array> struct ctype_char32 : std::ctype<char32_t> { ctype_char32(std::size_t refs = 0) : std::ctype<char32_t>(get_table(), false, refs) {} private: static const mask* get_table() { static std::array<mask, table_size> table; // Copy classification rules from ctype<wchar_t> const auto& wctype = std::use_facet<std::ctype<wchar_t>>(std::locale("en_US.utf8")); std::copy(wctype.table(), wctype.table() + std::ctype<wchar_t>::table_size, table.begin()); return table.data(); } }; bool foo(char32_t code_point) { // Attach our custom facet to the en_US.utf8 locale static std::locale loc(std::locale("en_US.utf8"), new ctype_char32); return std::isgraph(code_point, loc); } int main() { foo('B'); // No exception thrown now return 0; }
Portability note:
This first solution relies on wchar_t being 32-bit (true for Linux, but not Windows—where wchar_t is 16-bit). For cross-platform support, you’d need to handle UTF-16 conversion, but for Ubuntu systems, the first approach is straightforward and reliable.
内容的提问来源于stack exchange,提问作者spraff

