如何将含字节序列的字符串解析为字节序列并通过struct.unpack处理小端序数据?
struct.unpack Great question! The issue here boils down to how you're converting your string to bytes—using UTF-8 encoding is mangling the original byte values, which is why struct.unpack throws buffer size errors. Let's fix this and get your data parsed correctly without hardcoding the byte string directly.
Why Your Initial Approach Failed
When you run bytes(StrInstrument, 'utf-8'), Python tries to encode the string as UTF-8 characters. For example:
- The character
\xE0(which maps to the Latin-1 character à) gets encoded as two bytes:\xc3\xa0 - Similarly,
\xFFbecomes\xc3\xbf
This doubles the number of bytes in your data, making it impossible for struct.unpack to read the expected 8 bytes for 4 16-bit integers.
Fix 1: Use Latin-1 (ISO-8859-1) Encoding
Latin-1 maps every character from 0x00 to 0xFF directly to a single byte—perfect for preserving the original byte values in your string. Here's how to use it:
import struct StrInstrument = '\xE0\x31\xFF\xCF\xFF\xCA\xFF\xC4' # Convert string to bytes using Latin-1 to preserve original byte values byte_data = bytes(StrInstrument, 'latin-1') # Use '<HHHH' to explicitly specify little-endian (matches your expected output) iValues = struct.unpack('<HHHH', byte_data) print(iValues) # Output: (12768, 53247, 51967, 50431)
Fix 2: Read Data as Bytes Directly (If Possible)
If this string is coming from a file, serial port, or network stream, skip the string conversion entirely and read the data as bytes from the start. For example, reading from a binary file:
import struct with open('instrument_output.bin', 'rb') as f: # Read exactly 8 bytes (4 x 16-bit integers) byte_data = f.read(8) iValues = struct.unpack('<HHHH', byte_data) print(iValues)
Fix 3: Parse the String as a Byte Literal
If your string is a literal with escape sequences (like the one you provided), you can use ast.literal_eval to convert it directly to a bytes object without modifying the values:
import ast import struct StrInstrument = '\xE0\x31\xFF\xCF\xFF\xCA\xFF\xC4' # Wrap the string in b'' to treat it as a byte literal, then parse it byte_data = ast.literal_eval(f"b'{StrInstrument}'") iValues = struct.unpack('<HHHH', byte_data) print(iValues)
Key Takeaway
Always use an encoding that preserves 1:1 character-to-byte mapping (like Latin-1) when converting strings that represent raw binary data. UTF-8 is designed for text, not raw bytes, so it will alter your data unexpectedly.
内容的提问来源于stack exchange,提问作者martemis

