You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为UTF-8编码的单字节字符添加前导零?附C#字节转换实现代码

解决单字节字符添加前导零扩展的问题

Let's break this down clearly for you. First, I notice your current code uses Encoding.UTF32.GetBytes()—which encodes every character as a 4-byte value, so you shouldn't be seeing 1 or 2-byte characters here. That makes me think maybe you were using a different encoding (like UTF-8 or ASCII) before switching to UTF32, and now you want to ensure even single-byte characters get padded to 2 bytes with a leading zero.

最简单的解决方案:改用UTF-16 (Unicode)编码

If your goal is to have ASCII-range characters stored as 2-byte values with a leading 0x00 (and handle other characters appropriately), switching to UTF-16 is the most straightforward approach. In .NET, Encoding.Unicode corresponds to UTF-16 (little-endian), which does exactly this:

  • ASCII characters (0-127) become two bytes: 0x00 followed by the ASCII value
  • Most non-ASCII characters (like accented letters or Chinese/Japanese/Korean characters) are stored as two bytes
  • Rare supplementary characters use four bytes (but this is an edge case for most use cases)

Here's how to modify your code:

public byte[] ToByte() { 
    List<byte> result = new List<byte>(); 
    // First four bytes are for the Command 
    result.AddRange(BitConverter.GetBytes((int)cmdCommand)); 
    // Add the length of the name (character count)
    result.AddRange(BitConverter.GetBytes(strName?.Length ?? 0)); 
    // Length of the message (character count)
    result.AddRange(BitConverter.GetBytes(strMessage?.Length ?? 0)); 
    // Add the name using UTF-16
    if (strName != null) 
        result.AddRange(Encoding.Unicode.GetBytes(strName)); 
    // Add the message text using UTF-16
    if (strMessage != null) 
        result.AddRange(Encoding.Unicode.GetBytes(strMessage)); 
    return result.ToArray(); 
}

手动控制(如果需要强制所有字符为2字节)

If you need absolute control (e.g., even supplementary characters should be handled as two bytes, though this isn't standard), you can manually convert each char to a 2-byte array:

public byte[] ToByte() { 
    List<byte> result = new List<byte>(); 
    // First four bytes are for the Command 
    result.AddRange(BitConverter.GetBytes((int)cmdCommand)); 
    // Add the length of the name (character count)
    result.AddRange(BitConverter.GetBytes(strName?.Length ?? 0)); 
    // Length of the message (character count)
    result.AddRange(BitConverter.GetBytes(strMessage?.Length ?? 0)); 
    // Process name characters manually
    if (strName != null)
    {
        foreach (char c in strName)
        {
            // Convert char to 2-byte little-endian array
            byte[] charBytes = BitConverter.GetBytes(c);
            result.AddRange(charBytes);
        }
    }
    // Process message characters manually
    if (strMessage != null)
    {
        foreach (char c in strMessage)
        {
            byte[] charBytes = BitConverter.GetBytes(c);
            result.AddRange(charBytes);
        }
    }
    return result.ToArray(); 
}

关于长度存储的重要提示

Your current code stores the character count of the strings, not the encoded byte count. Make sure your decoding logic matches this: when reading the byte array back, you'll need to read the character count, then read count * 2 bytes (for UTF-16), then use Encoding.Unicode.GetString() to restore the original string. If you ever need to store the byte count instead, replace strName.Length with Encoding.Unicode.GetBytes(strName).Length.

内容的提问来源于stack exchange,提问作者ledmatrix

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 18:07:38