You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于glibc printf函数在nb_NO.utf8 locale下分组字符宽度计算错误的Bug报告及wchar_t字符串终止异常的技术问询

Hey there! Let's tackle your two issues step by step

First: Fixing the wchar_t string termination issue

Your code has a critical mistake in how you're defining the wide character: you're treating the UTF-8 byte sequence 0xe280af as a single integer value and casting it to wchar_t, which is completely wrong.

  • 0xe280af is three bytes of UTF-8 encoding that represents the Unicode code point U+200F (the Right-to-Left Mark). A wchar_t stores the actual Unicode code point, not the concatenated UTF-8 bytes. So the correct way to define this wide character is:
    wchar_t wc = 0x200F; // Alternatively, use the Unicode escape: L'\u200F'
    
  • When you use the invalid value 0xe280af in your string s2, it's not a valid Unicode code point. This causes functions like wcswidth to fail (returning -1) and makes printf("%ls") behave unpredictably—what looks like a non-terminating string is actually the result of invalid wide characters breaking the string parsing.

If you need to convert a UTF-8 byte sequence to a wide character properly (instead of hardcoding the code point), use the mbtowc function which respects your current locale:

char utf8_group_char[] = "\xe2\x80\xaf";
wchar_t wc;
// Convert the UTF-8 bytes to a wide character
mbtowc(&wc, utf8_group_char, sizeof(utf8_group_char) - 1);

Once you fix this, your s2 string will be valid, and both wcswidth and printf will handle it correctly.

Second: Confirming the glibc printf grouping character width bug

From your test output and code, this definitely looks like a bug in glibc's printf implementation when using the nb_NO.utf8 locale. Here's the breakdown:

  1. The symptom:

    • In nb_NO.utf8, printf("%'7d", 1234) outputs <1 234>. strlen(buf) returns 7 (telling us the grouping character is a 3-byte UTF-8 sequence), but your original invalid wide character code messed up the swprintf result—fix that first to get accurate width info.
    • Compare this to en_US.utf8, where the output is < 1,234>: both strlen and wcswidth return 7, which is correct—the format specifier %'7d asks for a total width of 7, so the left padding plus the number (with grouping) adds up properly.
    • In nb_NO.utf8, there's no left padding at all. This means printf is miscalculating the display width of the grouping character: it's treating the 3-byte UTF-8 sequence as taking up 3 width units instead of 1 (the actual display width of the character), so it thinks the output already meets the 7-width requirement without needing padding.
  2. Next steps to confirm:

    • First fix the wchar_t issue in your test code so swprintf generates a valid wide string. Then re-run the test to get an accurate wcswidth value (it should return 5 for 1 234, since the grouping character is 1 display unit).
    • With that fixed, you'll clearly see that printf isn't adding the required 2 spaces of left padding to reach the 7-width total. This confirms the bug in glibc's width calculation for locale-specific grouping characters.

You should report this bug to the glibc Bugzilla, including your corrected test code and side-by-side output from nb_NO.utf8 and en_US.utf8 locales to demonstrate the issue clearly.


内容的提问来源于stack exchange,提问作者oligofren

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 15:02:29