如何为UTF-8编码的单字节字符添加前导零?附C#字节转换实现代码
Let's break this down clearly for you. First, I notice your current code uses Encoding.UTF32.GetBytes()—which encodes every character as a 4-byte value, so you shouldn't be seeing 1 or 2-byte characters here. That makes me think maybe you were using a different encoding (like UTF-8 or ASCII) before switching to UTF32, and now you want to ensure even single-byte characters get padded to 2 bytes with a leading zero.
最简单的解决方案:改用UTF-16 (Unicode)编码
If your goal is to have ASCII-range characters stored as 2-byte values with a leading 0x00 (and handle other characters appropriately), switching to UTF-16 is the most straightforward approach. In .NET, Encoding.Unicode corresponds to UTF-16 (little-endian), which does exactly this:
- ASCII characters (0-127) become two bytes:
0x00followed by the ASCII value - Most non-ASCII characters (like accented letters or Chinese/Japanese/Korean characters) are stored as two bytes
- Rare supplementary characters use four bytes (but this is an edge case for most use cases)
Here's how to modify your code:
public byte[] ToByte() { List<byte> result = new List<byte>(); // First four bytes are for the Command result.AddRange(BitConverter.GetBytes((int)cmdCommand)); // Add the length of the name (character count) result.AddRange(BitConverter.GetBytes(strName?.Length ?? 0)); // Length of the message (character count) result.AddRange(BitConverter.GetBytes(strMessage?.Length ?? 0)); // Add the name using UTF-16 if (strName != null) result.AddRange(Encoding.Unicode.GetBytes(strName)); // Add the message text using UTF-16 if (strMessage != null) result.AddRange(Encoding.Unicode.GetBytes(strMessage)); return result.ToArray(); }
手动控制(如果需要强制所有字符为2字节)
If you need absolute control (e.g., even supplementary characters should be handled as two bytes, though this isn't standard), you can manually convert each char to a 2-byte array:
public byte[] ToByte() { List<byte> result = new List<byte>(); // First four bytes are for the Command result.AddRange(BitConverter.GetBytes((int)cmdCommand)); // Add the length of the name (character count) result.AddRange(BitConverter.GetBytes(strName?.Length ?? 0)); // Length of the message (character count) result.AddRange(BitConverter.GetBytes(strMessage?.Length ?? 0)); // Process name characters manually if (strName != null) { foreach (char c in strName) { // Convert char to 2-byte little-endian array byte[] charBytes = BitConverter.GetBytes(c); result.AddRange(charBytes); } } // Process message characters manually if (strMessage != null) { foreach (char c in strMessage) { byte[] charBytes = BitConverter.GetBytes(c); result.AddRange(charBytes); } } return result.ToArray(); }
关于长度存储的重要提示
Your current code stores the character count of the strings, not the encoded byte count. Make sure your decoding logic matches this: when reading the byte array back, you'll need to read the character count, then read count * 2 bytes (for UTF-16), then use Encoding.Unicode.GetString() to restore the original string. If you ever need to store the byte count instead, replace strName.Length with Encoding.Unicode.GetBytes(strName).Length.
内容的提问来源于stack exchange,提问作者ledmatrix

