You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Java中解码Python风格的UTF-8字节编码字符串

Java解码Python风格UTF-8字节字符串解决方案

核心思路

输入的文件内容是Python的字节字面量(如b'\nHola,\n\nAqu\xc3\xad est\xc3\xa1 la informaci\xc3\xb3n...'),本质是描述字节序列的字符串,而非原始字节数组。常规的getBytes()方法无法处理这类转义序列,必须先解析字符串中的转义规则,还原出真实的字节数组,再用UTF-8解码为正常文本。

实现代码

以下是纯Java内置API实现的解码方法,无需额外依赖:

import java.io.ByteArrayOutputStream;
import java.nio.charset.StandardCharsets;

public class PythonBytesDecoder {
    public static String decode(String pythonBytesLiteral) {
        // 移除Python字节字面量的标识:开头的b'/" 和结尾的'/""
        String content = pythonBytesLiteral.replaceAll("^b(['\"])|(['\"])$", "");
        ByteArrayOutputStream byteStream = new ByteArrayOutputStream();
        int strLength = content.length();
        int index = 0;

        while (index < strLength) {
            char currentChar = content.charAt(index);
            if (currentChar == '\\' && index + 1 < strLength) {
                char nextChar = content.charAt(index + 1);
                // 处理十六进制转义 \xXX
                if (nextChar == 'x' && index + 3 <= strLength) {
                    String hexSegment = content.substring(index + 2, index + 4);
                    try {
                        byte hexByte = (byte) Integer.parseInt(hexSegment, 16);
                        byteStream.write(hexByte);
                    } catch (NumberFormatException e) {
                        // 无效十六进制,直接保留原始字符
                        byteStream.write(currentChar);
                        byteStream.write(nextChar);
                        byteStream.write(hexSegment.getBytes());
                    }
                    index += 4;
                    continue;
                }
                // 处理普通转义字符
                switch (nextChar) {
                    case 'n':
                        byteStream.write('\n');
                        break;
                    case 't':
                        byteStream.write('\t');
                        break;
                    case 'r':
                        byteStream.write('\r');
                        break;
                    case '\\':
                        byteStream.write('\\');
                        break;
                    case '\'':
                        byteStream.write('\'');
                        break;
                    case '"':
                        byteStream.write('"');
                        break;
                    default:
                        // 未知转义,保留原始字符
                        byteStream.write(currentChar);
                        byteStream.write(nextChar);
                }
                index += 2;
            } else {
                // 普通字符直接写入
                byteStream.write(currentChar);
                index++;
            }
        }
        return new String(byteStream.toByteArray(), StandardCharsets.UTF_8);
    }

    public static void main(String[] args) {
        // 测试示例输入
        String input = "b'\\nHola,\\n\\nAqu\\xc3\\xad est\\xc3\\xa1 la informaci\\xc3\\xb3n...'";
        String decodedText = decode(input);
        System.out.println(decodedText);
    }
}

代码说明

  1. 清理输入:通过正则表达式移除Python字节字面量的b'/b"前缀和结尾引号,只保留核心的转义字符串。
  2. 解析转义序列:
    • 对\xXX格式的十六进制转义,将其解析为对应的字节并写入字节流。
    • 对\n、\t等普通转义字符,替换为对应的ASCII控制字符。
  3. 解码为UTF-8文本:将还原后的字节数组用UTF-8编码解码,得到包含多语言字符的正常文本。

内容的提问来源于stack exchange,提问作者Edgar Balderas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 12:43:37