如何将Hessian/Java编码的代理对二进制转Unicode字符?
UTF-16代理对解析实现(Rebol/Red版)
问题说明
- 待解析二进制数据:
#{EDA0BDEDB883}(Hessian/Java编码),目标解码结果为表情字符😃(对应Unicode码点U+1F603) - 核心问题:如何将上述二进制数据转换为UTF-16代理对
\ud83d和\ude03 - 需求:将下方Python3的无依赖二进制解析逻辑,重写为Rebol或Red-lang代码
原Python3实现代码
def _decode_surrogate_pair(c1, c2): """ Python 3不再自动解码代理对,需手动处理 """ # print(c1.encode('utf-8')) # \ud83d # print(c2.encode('utf-8')) # \ude03 if not('\uD800' <= c1 <= '\uDBFF') or not ('\uDC00' <= c2 <= '\uDFFF'): raise Exception("无效的UTF-16代理对") code = 0x10000 code += (ord(c1) & 0x03FF) << 10 code += (ord(c2) & 0x03FF) return chr(code) def _decode_byte_array(bytes): s = '' while(len(bytes)): b, bytes = bytes[0], bytes[1:] c = b.decode('utf-8', 'surrogatepass') if '\uD800' <= c <= '\uDBFF': b, bytes = bytes[0], bytes[1:] c2 = b.decode('utf-8', 'surrogatepass') c = _decode_surrogate_pair(c, c2) s += c return s # 测试用二进制数据,对应#{EDA0BDEDB883} bytes = [b'\xed\xa0\xbd', b'\xed\xb8\x83'] print(_decode_byte_array(bytes))
Java参考代码示例
import java.util.Arrays; import java.nio.charset.StandardCharsets; public class SurrogatePairDemo { public static void main(String[] args) throws Exception { // 目标字符:"😃",对应UTF-16代理对"\uD83D\uDE03" final byte[] bytes1 = "😃".getBytes(StandardCharsets.UTF_16); // 输出结果:[-2, -1, -40, 61, -34, 3] // 对应十六进制:#{FE FF D8 3D DE 03}(UTF-16大端带BOM) System.out.println(Arrays.toString(bytes1)); } }
Rebol/Red实现代码
核心逻辑说明
- 按UTF-8规则拆分二进制数据,解码为单个字符(保留代理字符)
- 检测高代理字符(范围
U+D800到U+DBFF),若存在则读取后续的低代理字符(U+DC00到U+DFFF) - 通过公式计算完整Unicode码点:
code = 0x10000 + (高代理字符 & 0x03FF) << 10 + (低代理字符 & 0x03FF) - 将码点转换为对应字符
Rebol版本代码
decode-surrogate-pair: func [c1 c2 [char!] /local code] [ if not (c1 >= to char! 0xD800 and c1 <= to char! 0xDBFF) or not (c2 >= to char! 0xDC00 and c2 <= to char! 0xDFFF) [ fail "无效的UTF-16代理对" ] code: 0x10000 code: code + ((to integer! c1) & 0x03FF) << 10 code: code + ((to integer! c2) & 0x03FF) return to char! code ] decode-byte-array: func [bytes [block! binary!] /local s b c c2 temp len first-byte] [ s: copy "" either block? bytes [ while [not empty? bytes] [ b: first bytes bytes: next bytes c: to string! b if (to integer! c) >= 0xD800 and (to integer! c) <= 0xDBFF [ if empty? bytes [fail "不完整的代理对"] b: first bytes bytes: next bytes c2: to string! b c: decode-surrogate-pair c c2 ] append s c ] ] [ temp: copy bytes while [not empty? temp] [ first-byte: pick temp 1 if first-byte < 0x80 [len: 1] else if first-byte < 0xE0 [len: 2] else if first-byte < 0xF0 [len: 3] else [len: 4] b: copy/part temp len temp: skip temp len c: to string! b if (to integer! c) >= 0xD800 and (to integer! c) <= 0xDBFF [ if empty? temp [fail "不完整的代理对"] first-byte: pick temp 1 if first-byte < 0x80 [len: 1] else if first-byte < 0xE0 [len: 2] else if first-byte < 0xF0 [len: 3] else [len: 4] b2: copy/part temp len temp: skip temp len c2: to string! b2 c: decode-surrogate-pair c c2 ] append s c ] ] return s ] ; 测试用例1:分块二进制数据 test-bytes: [#{EDA0BD} #{EDB883}] print decode-byte-array test-bytes ; 输出😃 ; 测试用例2:完整二进制数据 test-binary: #{EDA0BDEDB883} print decode-byte-array test-binary ; 输出😃
Red版本代码
decode-surrogate-pair: function [c1 c2 [char!] /local code][ if not (c1 >= to char! 0xD800 and c1 <= to char! 0xDBFF) or not (c2 >= to char! 0xDC00 and c2 <= to char! 0xDFFF)[ throw "无效的UTF-16代理对" ] code: 0x10000 code: code + ((to integer! c1) & 0x03FF) << 10 code: code + ((to integer! c2) & 0x03FF) return to char! code ] decode-byte-array: function [bytes [block! binary!] /local s b c c2 temp len first-byte][ s: copy "" either block? bytes [ while [not empty? bytes][ b: first bytes bytes: next bytes c: to string! b if (to integer! c) >= 0xD800 and (to integer! c) <= 0xDBFF [ if empty? bytes [throw "不完整的代理对"] b: first bytes bytes: next bytes c2: to string! b c: decode-surrogate-pair c c2 ] append s c ] ][ temp: copy bytes while [not empty? temp][ first-byte: pick temp 1 if first-byte < 0x80 [len: 1] else if first-byte < 0xE0 [len: 2] else if first-byte < 0xF0 [len: 3] else [len: 4] b: copy/part temp len temp: skip temp len c: to string! b if (to integer! c) >= 0xD800 and (to integer! c) <= 0xDBFF [ if empty? temp [throw "不完整的代理对"] first-byte: pick temp 1 if first-byte < 0x80 [len: 1] else if first-byte < 0xE0 [len: 2] else if first-byte < 0xF0 [len: 3] else [len: 4] b2: copy/part temp len temp: skip temp len c2: to string! b2 c: decode-surrogate-pair c c2 ] append s c ] ] return s ] ; 测试用例 test-bytes: [#{EDA0BD} #{EDB883}] print decode-byte-array test-bytes test-binary: #{EDA0BDEDB883} print decode-byte-array test-binary
内容的提问来源于stack exchange,提问作者tianwen
相关产品推荐
相关产品推荐

