You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Swift 5中如何解码"\xc3\xa6"这类UTF-8字面量?

在Swift 5中解码字符串中的UTF-8字面量

你遇到的问题核心是:你拿到的字符串里的\xc3\xa6并不是真正的UTF-8字节,而是由**普通字符(反斜杠、x、十六进制数字)**组成的转义序列。直接调用s.utf8只会获取这些普通字符本身的UTF-8编码,自然无法得到你想要的æ字符。

要解决这个问题,我们需要先把这些转义的十六进制序列转换成对应的UInt8字节,再用这些字节构建真正的UTF-8字符串。下面是一个实用的实现方案:

实现解码函数

func decodeEscapedUTF8(_ input: String) -> String? {
    var byteBuffer = [UInt8]()
    var currentIndex = input.startIndex
    
    while currentIndex < input.endIndex {
        // 检查是否遇到转义符 \
        if input[currentIndex] == "\\" && currentIndex < input.index(before: input.endIndex) {
            let nextChar = input[input.index(after: currentIndex)]
            // 匹配 \x 开头的十六进制转义序列
            if nextChar == "x" && input.index(after: currentIndex) < input.index(before: input.endIndex) {
                // 提取两位十六进制字符
                let hexStart = input.index(currentIndex, offsetBy: 2)
                let hexEnd = input.index(hexStart, offsetBy: 2)
                guard hexEnd <= input.endIndex else {
                    return nil // 转义序列不完整,返回nil
                }
                let hexStr = String(input[hexStart..<hexEnd])
                // 把十六进制字符串转成UInt8字节
                if let byte = UInt8(hexStr, radix: 16) {
                    byteBuffer.append(byte)
                    currentIndex = hexEnd
                    continue
                }
            }
            // 如果不是 \x 转义,直接添加反斜杠本身的UTF-8字节
            byteBuffer.append(contentsOf: "\\".utf8)
            currentIndex = input.index(after: currentIndex)
        } else {
            // 普通字符,直接添加其UTF-8字节
            let charStr = String(input[currentIndex])
            byteBuffer.append(contentsOf: charStr.utf8)
            currentIndex = input.index(after: currentIndex)
        }
    }
    
    // 用转换后的字节数组构建UTF-8字符串
    return String(bytes: byteBuffer, encoding: .utf8)
}

测试使用

let encodedSSID = "\\xc3\\xa6"
if let decodedSSID = decodeEscapedUTF8(encodedSSID) {
    print(decodedSSID) // 输出:æ
}

原理说明

  1. 转义序列识别:函数会遍历输入字符串,识别\x开头的十六进制转义序列,将其转换成对应的UInt8字节。
  2. 字节构建与解码:收集所有转换后的字节(包括普通字符的UTF-8字节),最后用String(bytes:encoding:)方法将字节数组解码成真正的UTF-8字符串。
  3. 容错处理:如果遇到不完整的转义序列(比如\x后面不足两位十六进制数),函数会返回nil,你可以根据需求调整容错逻辑。

内容的提问来源于stack exchange,提问作者YoungChul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 03:12:40