Java中按字节长度限制分割字符串的技术实现问询
Alright, let's break down how to solve this problem properly. The biggest gotcha here is handling multi-byte Unicode characters (like Chinese characters, emojis, or other non-ASCII symbols) — if you try to split a string directly by byte count, you risk cutting a multi-byte character in half, which results in corrupted garbage when you decode it later.
Below is a step-by-step implementation approach, followed by a complete Java solution and test cases to cover all edge scenarios.
The core idea is to iterate through the string by code points (not individual chars, since some Unicode characters are made of two chars called surrogate pairs) and track the cumulative byte length of the current segment. Here's the play-by-play:
- Validate input parameters to avoid edge case crashes.
- Use a
Charsetobject to handle encoding safely (avoids deprecatedString.getBytes(String)methods). - For each code point, calculate its byte length in the target encoding before adding it to the current segment.
- If adding the next code point would exceed the
maxSizelimit, finalize the current segment, start a new one, and add the code point to it. - Handle edge cases like empty input,
maxSizesmaller than a single character's byte length, and strings that fit perfectly into one segment.
import java.nio.charset.Charset; import java.nio.charset.IllegalCharsetNameException; import java.nio.charset.UnsupportedCharsetException; import java.util.ArrayList; import java.util.List; public class StringByteSplitter { public static String[] splitByByteSize(String input, int maxSize, String encoding) { // Validate input parameters if (input == null) { throw new IllegalArgumentException("Input string cannot be null"); } if (maxSize <= 0) { throw new IllegalArgumentException("Max size must be a positive integer"); } if (encoding == null || encoding.isBlank()) { throw new IllegalArgumentException("Encoding cannot be null or blank"); } Charset charset; try { charset = Charset.forName(encoding); } catch (IllegalCharsetNameException | UnsupportedCharsetException e) { throw new IllegalArgumentException("Unsupported encoding: " + encoding, e); } if (input.isEmpty()) { return new String[0]; } List<String> segments = new ArrayList<>(); StringBuilder currentSegment = new StringBuilder(); int currentByteLength = 0; int length = input.length(); int i = 0; while (i < length) { int codePoint = input.codePointAt(i); // Convert code point to its string representation and get byte length String codePointStr = new String(Character.toChars(codePoint)); int codePointByteLength = codePointStr.getBytes(charset).length; // Check if this code point alone exceeds max size if (codePointByteLength > maxSize) { throw new IllegalArgumentException( String.format("Single character exceeds max size: %d bytes (max allowed: %d)", codePointByteLength, maxSize) ); } // Check if adding this code point would exceed max size if (currentByteLength + codePointByteLength > maxSize) { // Finalize current segment segments.add(currentSegment.toString()); currentSegment.setLength(0); currentByteLength = 0; } // Add the code point to current segment currentSegment.append(codePointStr); currentByteLength += codePointByteLength; // Move index past this code point (handles surrogate pairs) i += Character.charCount(codePoint); } // Add the last segment if (currentSegment.length() > 0) { segments.add(currentSegment.toString()); } return segments.toArray(new String[0]); } }
Let's use JUnit 5 to test all critical scenarios:
import org.junit.jupiter.api.Test; import static org.junit.jupiter.api.Assertions.*; public class StringByteSplitterTest { @Test void testAsciiString() { String input = "Hello World! This is an ASCII string."; String[] result = StringByteSplitter.splitByByteSize(input, 10, "UTF-8"); // Verify each segment's byte length is <=10 for (String segment : result) { assertTrue(segment.getBytes().length <= 10); } // Verify concatenation matches original assertEquals(input, String.join("", result)); } @Test void testMultiByteChinese() { String input = "我爱Java编程,这是一段中文测试。"; String[] result = StringByteSplitter.splitByByteSize(input, 9, "UTF-8"); // Each Chinese char is 3 bytes in UTF-8, so max 3 chars per segment (3*3=9) for (String segment : result) { assertTrue(segment.getBytes().length <=9); } assertEquals(input, String.join("", result)); } @Test void testSurrogatePairEmoji() { String input = "Hello 😀👋🌍! Let's test emojis."; String[] result = StringByteSplitter.splitByByteSize(input, 10, "UTF-8"); // Emojis are 4 bytes each in UTF-8 for (String segment : result) { assertTrue(segment.getBytes().length <=10); } assertEquals(input, String.join("", result)); } @Test void testMaxSizeExactlyMatches() { String input = "ABCDEFGHIJ"; // 10 ASCII chars = 10 bytes String[] result = StringByteSplitter.splitByByteSize(input, 10, "UTF-8"); assertEquals(1, result.length); assertEquals(input, result[0]); } @Test void testSingleCharacterExceedsMaxSize() { String input = "😀"; // 4 bytes in UTF-8 assertThrows(IllegalArgumentException.class, () -> StringByteSplitter.splitByByteSize(input, 3, "UTF-8")); } @Test void testEmptyInput() { String[] result = StringByteSplitter.splitByByteSize("", 5, "UTF-8"); assertEquals(0, result.length); } }
- Surrogate Pair Handling: Using
codePointAtandCharacter.charCountensures we don't split emojis or other supplementary characters (which use twochars) into invalid parts. - Encoding Safety: Using
Charset.forNameinstead of the deprecatedString.getBytes(String)method gives us proper exception handling for unsupported encodings. - Validation: We explicitly check for invalid inputs and impossible scenarios (like a single character exceeding
maxSize) to avoid silent failures.
内容的提问来源于stack exchange,提问作者Hyeonseo Yang

