You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java中按字节长度限制分割字符串的技术实现问询

Alright, let's break down how to solve this problem properly. The biggest gotcha here is handling multi-byte Unicode characters (like Chinese characters, emojis, or other non-ASCII symbols) — if you try to split a string directly by byte count, you risk cutting a multi-byte character in half, which results in corrupted garbage when you decode it later.

Below is a step-by-step implementation approach, followed by a complete Java solution and test cases to cover all edge scenarios.

Implementation Approach

The core idea is to iterate through the string by code points (not individual chars, since some Unicode characters are made of two chars called surrogate pairs) and track the cumulative byte length of the current segment. Here's the play-by-play:

  • Validate input parameters to avoid edge case crashes.
  • Use a Charset object to handle encoding safely (avoids deprecated String.getBytes(String) methods).
  • For each code point, calculate its byte length in the target encoding before adding it to the current segment.
  • If adding the next code point would exceed the maxSize limit, finalize the current segment, start a new one, and add the code point to it.
  • Handle edge cases like empty input, maxSize smaller than a single character's byte length, and strings that fit perfectly into one segment.
Java Implementation
import java.nio.charset.Charset;
import java.nio.charset.IllegalCharsetNameException;
import java.nio.charset.UnsupportedCharsetException;
import java.util.ArrayList;
import java.util.List;

public class StringByteSplitter {

    public static String[] splitByByteSize(String input, int maxSize, String encoding) {
        // Validate input parameters
        if (input == null) {
            throw new IllegalArgumentException("Input string cannot be null");
        }
        if (maxSize <= 0) {
            throw new IllegalArgumentException("Max size must be a positive integer");
        }
        if (encoding == null || encoding.isBlank()) {
            throw new IllegalArgumentException("Encoding cannot be null or blank");
        }

        Charset charset;
        try {
            charset = Charset.forName(encoding);
        } catch (IllegalCharsetNameException | UnsupportedCharsetException e) {
            throw new IllegalArgumentException("Unsupported encoding: " + encoding, e);
        }

        if (input.isEmpty()) {
            return new String[0];
        }

        List<String> segments = new ArrayList<>();
        StringBuilder currentSegment = new StringBuilder();
        int currentByteLength = 0;

        int length = input.length();
        int i = 0;
        while (i < length) {
            int codePoint = input.codePointAt(i);
            // Convert code point to its string representation and get byte length
            String codePointStr = new String(Character.toChars(codePoint));
            int codePointByteLength = codePointStr.getBytes(charset).length;

            // Check if this code point alone exceeds max size
            if (codePointByteLength > maxSize) {
                throw new IllegalArgumentException(
                    String.format("Single character exceeds max size: %d bytes (max allowed: %d)",
                        codePointByteLength, maxSize)
                );
            }

            // Check if adding this code point would exceed max size
            if (currentByteLength + codePointByteLength > maxSize) {
                // Finalize current segment
                segments.add(currentSegment.toString());
                currentSegment.setLength(0);
                currentByteLength = 0;
            }

            // Add the code point to current segment
            currentSegment.append(codePointStr);
            currentByteLength += codePointByteLength;

            // Move index past this code point (handles surrogate pairs)
            i += Character.charCount(codePoint);
        }

        // Add the last segment
        if (currentSegment.length() > 0) {
            segments.add(currentSegment.toString());
        }

        return segments.toArray(new String[0]);
    }
}
Test Cases

Let's use JUnit 5 to test all critical scenarios:

import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;

public class StringByteSplitterTest {

    @Test
    void testAsciiString() {
        String input = "Hello World! This is an ASCII string.";
        String[] result = StringByteSplitter.splitByByteSize(input, 10, "UTF-8");
        // Verify each segment's byte length is <=10
        for (String segment : result) {
            assertTrue(segment.getBytes().length <= 10);
        }
        // Verify concatenation matches original
        assertEquals(input, String.join("", result));
    }

    @Test
    void testMultiByteChinese() {
        String input = "我爱Java编程,这是一段中文测试。";
        String[] result = StringByteSplitter.splitByByteSize(input, 9, "UTF-8");
        // Each Chinese char is 3 bytes in UTF-8, so max 3 chars per segment (3*3=9)
        for (String segment : result) {
            assertTrue(segment.getBytes().length <=9);
        }
        assertEquals(input, String.join("", result));
    }

    @Test
    void testSurrogatePairEmoji() {
        String input = "Hello 😀👋🌍! Let's test emojis.";
        String[] result = StringByteSplitter.splitByByteSize(input, 10, "UTF-8");
        // Emojis are 4 bytes each in UTF-8
        for (String segment : result) {
            assertTrue(segment.getBytes().length <=10);
        }
        assertEquals(input, String.join("", result));
    }

    @Test
    void testMaxSizeExactlyMatches() {
        String input = "ABCDEFGHIJ"; // 10 ASCII chars = 10 bytes
        String[] result = StringByteSplitter.splitByByteSize(input, 10, "UTF-8");
        assertEquals(1, result.length);
        assertEquals(input, result[0]);
    }

    @Test
    void testSingleCharacterExceedsMaxSize() {
        String input = "😀"; // 4 bytes in UTF-8
        assertThrows(IllegalArgumentException.class,
            () -> StringByteSplitter.splitByByteSize(input, 3, "UTF-8"));
    }

    @Test
    void testEmptyInput() {
        String[] result = StringByteSplitter.splitByByteSize("", 5, "UTF-8");
        assertEquals(0, result.length);
    }
}
Key Notes
  • Surrogate Pair Handling: Using codePointAt and Character.charCount ensures we don't split emojis or other supplementary characters (which use two chars) into invalid parts.
  • Encoding Safety: Using Charset.forName instead of the deprecated String.getBytes(String) method gives us proper exception handling for unsupported encodings.
  • Validation: We explicitly check for invalid inputs and impossible scenarios (like a single character exceeding maxSize) to avoid silent failures.

内容的提问来源于stack exchange,提问作者Hyeonseo Yang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:07:36