You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Hessian/Java编码的代理对二进制转Unicode字符?

UTF-16代理对解析实现(Rebol/Red版)

问题说明

  • 待解析二进制数据:#{EDA0BDEDB883}(Hessian/Java编码),目标解码结果为表情字符😃(对应Unicode码点U+1F603)
  • 核心问题:如何将上述二进制数据转换为UTF-16代理对\ud83d和\ude03
  • 需求:将下方Python3的无依赖二进制解析逻辑,重写为Rebol或Red-lang代码

原Python3实现代码

def _decode_surrogate_pair(c1, c2):
    """
    Python 3不再自动解码代理对,需手动处理
    """
    # print(c1.encode('utf-8')) # \ud83d
    # print(c2.encode('utf-8')) # \ude03
    if not('\uD800' <= c1 <= '\uDBFF') or not ('\uDC00' <= c2 <= '\uDFFF'):
        raise Exception("无效的UTF-16代理对")
    code = 0x10000
    code += (ord(c1) & 0x03FF) << 10
    code += (ord(c2) & 0x03FF)
    return chr(code)

def _decode_byte_array(bytes):
    s = ''
    while(len(bytes)):
        b, bytes = bytes[0], bytes[1:]
   
        c = b.decode('utf-8', 'surrogatepass')
        if '\uD800' <= c <= '\uDBFF':
            b, bytes = bytes[0], bytes[1:]
            c2 = b.decode('utf-8', 'surrogatepass')
            c = _decode_surrogate_pair(c, c2)
        s += c
    return s

# 测试用二进制数据,对应#{EDA0BDEDB883}
bytes = [b'\xed\xa0\xbd', b'\xed\xb8\x83']

print(_decode_byte_array(bytes))

Java参考代码示例

import java.util.Arrays;
import java.nio.charset.StandardCharsets;

public class SurrogatePairDemo {
    public static void main(String[] args) throws Exception {
        // 目标字符:"😃",对应UTF-16代理对"\uD83D\uDE03"
        final byte[] bytes1 = "😃".getBytes(StandardCharsets.UTF_16);
        // 输出结果:[-2, -1, -40, 61, -34, 3]
        // 对应十六进制:#{FE FF D8 3D DE 03}(UTF-16大端带BOM)
        System.out.println(Arrays.toString(bytes1));
    }
}

Rebol/Red实现代码

核心逻辑说明

  1. 按UTF-8规则拆分二进制数据,解码为单个字符(保留代理字符)
  2. 检测高代理字符(范围U+D800到U+DBFF),若存在则读取后续的低代理字符(U+DC00到U+DFFF)
  3. 通过公式计算完整Unicode码点:code = 0x10000 + (高代理字符 & 0x03FF) << 10 + (低代理字符 & 0x03FF)
  4. 将码点转换为对应字符

Rebol版本代码

decode-surrogate-pair: func [c1 c2 [char!] /local code] [
    if not (c1 >= to char! 0xD800 and c1 <= to char! 0xDBFF) or not (c2 >= to char! 0xDC00 and c2 <= to char! 0xDFFF) [
        fail "无效的UTF-16代理对"
    ]
    code: 0x10000
    code: code + ((to integer! c1) & 0x03FF) << 10
    code: code + ((to integer! c2) & 0x03FF)
    return to char! code
]

decode-byte-array: func [bytes [block! binary!] /local s b c c2 temp len first-byte] [
    s: copy ""
    either block? bytes [
        while [not empty? bytes] [
            b: first bytes
            bytes: next bytes
            c: to string! b
            if (to integer! c) >= 0xD800 and (to integer! c) <= 0xDBFF [
                if empty? bytes [fail "不完整的代理对"]
                b: first bytes
                bytes: next bytes
                c2: to string! b
                c: decode-surrogate-pair c c2
            ]
            append s c
        ]
    ] [
        temp: copy bytes
        while [not empty? temp] [
            first-byte: pick temp 1
            if first-byte < 0x80 [len: 1]
            else if first-byte < 0xE0 [len: 2]
            else if first-byte < 0xF0 [len: 3]
            else [len: 4]
            b: copy/part temp len
            temp: skip temp len
            c: to string! b
            if (to integer! c) >= 0xD800 and (to integer! c) <= 0xDBFF [
                if empty? temp [fail "不完整的代理对"]
                first-byte: pick temp 1
                if first-byte < 0x80 [len: 1]
                else if first-byte < 0xE0 [len: 2]
                else if first-byte < 0xF0 [len: 3]
                else [len: 4]
                b2: copy/part temp len
                temp: skip temp len
                c2: to string! b2
                c: decode-surrogate-pair c c2
            ]
            append s c
        ]
    ]
    return s
]

; 测试用例1:分块二进制数据
test-bytes: [#{EDA0BD} #{EDB883}]
print decode-byte-array test-bytes ; 输出😃

; 测试用例2:完整二进制数据
test-binary: #{EDA0BDEDB883}
print decode-byte-array test-binary ; 输出😃

Red版本代码

decode-surrogate-pair: function [c1 c2 [char!] /local code][
    if not (c1 >= to char! 0xD800 and c1 <= to char! 0xDBFF) or not (c2 >= to char! 0xDC00 and c2 <= to char! 0xDFFF)[
        throw "无效的UTF-16代理对"
    ]
    code: 0x10000
    code: code + ((to integer! c1) & 0x03FF) << 10
    code: code + ((to integer! c2) & 0x03FF)
    return to char! code
]

decode-byte-array: function [bytes [block! binary!] /local s b c c2 temp len first-byte][
    s: copy ""
    either block? bytes [
        while [not empty? bytes][
            b: first bytes
            bytes: next bytes
            c: to string! b
            if (to integer! c) >= 0xD800 and (to integer! c) <= 0xDBFF [
                if empty? bytes [throw "不完整的代理对"]
                b: first bytes
                bytes: next bytes
                c2: to string! b
                c: decode-surrogate-pair c c2
            ]
            append s c
        ]
    ][
        temp: copy bytes
        while [not empty? temp][
            first-byte: pick temp 1
            if first-byte < 0x80 [len: 1]
            else if first-byte < 0xE0 [len: 2]
            else if first-byte < 0xF0 [len: 3]
            else [len: 4]
            b: copy/part temp len
            temp: skip temp len
            c: to string! b
            if (to integer! c) >= 0xD800 and (to integer! c) <= 0xDBFF [
                if empty? temp [throw "不完整的代理对"]
                first-byte: pick temp 1
                if first-byte < 0x80 [len: 1]
                else if first-byte < 0xE0 [len: 2]
                else if first-byte < 0xF0 [len: 3]
                else [len: 4]
                b2: copy/part temp len
                temp: skip temp len
                c2: to string! b2
                c: decode-surrogate-pair c c2
            ]
            append s c
        ]
    ]
    return s
]

; 测试用例
test-bytes: [#{EDA0BD} #{EDB883}]
print decode-byte-array test-bytes

test-binary: #{EDA0BDEDB883}
print decode-byte-array test-binary

内容的提问来源于stack exchange,提问作者tianwen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 03:30:59