如何在字节数组中分隔字符串?边缘场景处理实践咨询
Hey there! Switching from JSON to ArrayBuffer for string storage definitely brings up those tricky edge-case headaches with delimiters—let’s walk through practical solutions to tackle them.
1. Escape the Delimiter (For Quotes or Custom Sequences)
If you want to stick with a delimiter like double quotes or a custom sequence, the classic fix is escaping any occurrence of the delimiter inside your strings.
Example with Double Quotes:
- When writing a string like
"str"ing", replace internal double quotes with an escaped version (e.g.,\"). Your stored string becomes"str\"ing". - When reading back, split on unescaped double quotes, then reverse the escape (replace
\"with") to recover the original string.
Example with Custom Delimiters:
If you’re using --MYDELIMITER-- as a separator, scan each string for this sequence and escape a part of it—say, replace --MYDELIMITER-- with --MYDELIMITER\- inside the string. When parsing, first convert the escaped sequences back to the original, then split on the unescaped --MYDELIMITER--.
Pros: Familiar if you’re used to string escaping patterns.
Cons: Adds parsing complexity, and there’s always a tiny risk you miss an edge case (like nested escapes).
2. Length-Prefixing (The Most Robust Solution)
This is my go-to recommendation for ArrayBuffer string storage—it completely avoids delimiter conflicts by storing metadata about each string before the string itself.
Here’s how it works:
- When writing:
- Convert each string to a UTF-8 byte array.
- Write a fixed-size length value (e.g., a 4-byte
Uint32integer) representing the number of bytes in the UTF-8 array. - Write the UTF-8 bytes right after the length prefix.
- When reading:
- Read the 4-byte length value first to know how many bytes to read next.
- Read exactly that number of bytes, convert them back to a string, and repeat for the next entry.
Example Code Snippet (JavaScript):
// Writing strings to ArrayBuffer function writeStringsToBuffer(strings) { const encoder = new TextEncoder(); // Calculate total buffer size: 4 bytes per length + sum of UTF-8 byte lengths const totalSize = strings.reduce((acc, str) => acc + 4 + encoder.encode(str).length, 0); const buffer = new ArrayBuffer(totalSize); const view = new DataView(buffer); let offset = 0; strings.forEach(str => { const bytes = encoder.encode(str); // Write length (4 bytes, little-endian) view.setUint32(offset, bytes.length, true); offset += 4; // Write string bytes new Uint8Array(buffer, offset, bytes.length).set(bytes); offset += bytes.length; }); return buffer; } // Reading strings from ArrayBuffer function readStringsFromBuffer(buffer) { const decoder = new TextDecoder(); const view = new DataView(buffer); const strings = []; let offset = 0; while (offset < buffer.byteLength) { // Read length const length = view.getUint32(offset, true); offset += 4; // Read string bytes and decode const bytes = new Uint8Array(buffer, offset, length); strings.push(decoder.decode(bytes)); offset += length; } return strings; }
Pros: No delimiter conflicts ever—works with any string content, no matter how unusual. Parsing is straightforward and error-resistant.
Cons: Adds a small overhead (4 bytes per string), but this is negligible for almost all use cases.
3. Use a "Magic" Delimiter (Last Resort)
If you really want to avoid length prefixes, pick a delimiter sequence that’s extremely unlikely to appear in your strings—like a combination of rare Unicode characters (e.g., \u0001\u0002\u0003) or a very long random sequence (e.g., XYZ123DELIMXYZ123).
But even then, you should still add an escape mechanism for the rare case the sequence does show up. This is less reliable than length prefixing, but it’s an option if you have strict size constraints.
In most scenarios, length-prefixing is the way to go—it eliminates all ambiguity of delimiters and handles every edge case you can throw at it.
内容的提问来源于stack exchange,提问作者Lance Pollard

