关于glibc printf函数在nb_NO.utf8 locale下分组字符宽度计算错误的Bug报告及wchar_t字符串终止异常的技术问询
First: Fixing the wchar_t string termination issue
Your code has a critical mistake in how you're defining the wide character: you're treating the UTF-8 byte sequence 0xe280af as a single integer value and casting it to wchar_t, which is completely wrong.
0xe280afis three bytes of UTF-8 encoding that represents the Unicode code pointU+200F(the Right-to-Left Mark). Awchar_tstores the actual Unicode code point, not the concatenated UTF-8 bytes. So the correct way to define this wide character is:wchar_t wc = 0x200F; // Alternatively, use the Unicode escape: L'\u200F'- When you use the invalid value
0xe280afin your strings2, it's not a valid Unicode code point. This causes functions likewcswidthto fail (returning -1) and makesprintf("%ls")behave unpredictably—what looks like a non-terminating string is actually the result of invalid wide characters breaking the string parsing.
If you need to convert a UTF-8 byte sequence to a wide character properly (instead of hardcoding the code point), use the mbtowc function which respects your current locale:
char utf8_group_char[] = "\xe2\x80\xaf"; wchar_t wc; // Convert the UTF-8 bytes to a wide character mbtowc(&wc, utf8_group_char, sizeof(utf8_group_char) - 1);
Once you fix this, your s2 string will be valid, and both wcswidth and printf will handle it correctly.
Second: Confirming the glibc printf grouping character width bug
From your test output and code, this definitely looks like a bug in glibc's printf implementation when using the nb_NO.utf8 locale. Here's the breakdown:
The symptom:
- In
nb_NO.utf8,printf("%'7d", 1234)outputs<1 234>.strlen(buf)returns 7 (telling us the grouping character is a 3-byte UTF-8 sequence), but your original invalid wide character code messed up theswprintfresult—fix that first to get accurate width info. - Compare this to
en_US.utf8, where the output is< 1,234>: bothstrlenandwcswidthreturn 7, which is correct—the format specifier%'7dasks for a total width of 7, so the left padding plus the number (with grouping) adds up properly. - In
nb_NO.utf8, there's no left padding at all. This meansprintfis miscalculating the display width of the grouping character: it's treating the 3-byte UTF-8 sequence as taking up 3 width units instead of 1 (the actual display width of the character), so it thinks the output already meets the 7-width requirement without needing padding.
- In
Next steps to confirm:
- First fix the
wchar_tissue in your test code soswprintfgenerates a valid wide string. Then re-run the test to get an accuratewcswidthvalue (it should return 5 for1 234, since the grouping character is 1 display unit). - With that fixed, you'll clearly see that
printfisn't adding the required 2 spaces of left padding to reach the 7-width total. This confirms the bug in glibc's width calculation for locale-specific grouping characters.
- First fix the
You should report this bug to the glibc Bugzilla, including your corrected test code and side-by-side output from nb_NO.utf8 and en_US.utf8 locales to demonstrate the issue clearly.
内容的提问来源于stack exchange,提问作者oligofren

