Swift 5中如何解码"\xc3\xa6"这类UTF-8字面量?
在Swift 5中解码字符串中的UTF-8字面量
你遇到的问题核心是:你拿到的字符串里的\xc3\xa6并不是真正的UTF-8字节,而是由**普通字符(反斜杠、x、十六进制数字)**组成的转义序列。直接调用s.utf8只会获取这些普通字符本身的UTF-8编码,自然无法得到你想要的æ字符。
要解决这个问题,我们需要先把这些转义的十六进制序列转换成对应的UInt8字节,再用这些字节构建真正的UTF-8字符串。下面是一个实用的实现方案:
实现解码函数
func decodeEscapedUTF8(_ input: String) -> String? { var byteBuffer = [UInt8]() var currentIndex = input.startIndex while currentIndex < input.endIndex { // 检查是否遇到转义符 \ if input[currentIndex] == "\\" && currentIndex < input.index(before: input.endIndex) { let nextChar = input[input.index(after: currentIndex)] // 匹配 \x 开头的十六进制转义序列 if nextChar == "x" && input.index(after: currentIndex) < input.index(before: input.endIndex) { // 提取两位十六进制字符 let hexStart = input.index(currentIndex, offsetBy: 2) let hexEnd = input.index(hexStart, offsetBy: 2) guard hexEnd <= input.endIndex else { return nil // 转义序列不完整,返回nil } let hexStr = String(input[hexStart..<hexEnd]) // 把十六进制字符串转成UInt8字节 if let byte = UInt8(hexStr, radix: 16) { byteBuffer.append(byte) currentIndex = hexEnd continue } } // 如果不是 \x 转义,直接添加反斜杠本身的UTF-8字节 byteBuffer.append(contentsOf: "\\".utf8) currentIndex = input.index(after: currentIndex) } else { // 普通字符,直接添加其UTF-8字节 let charStr = String(input[currentIndex]) byteBuffer.append(contentsOf: charStr.utf8) currentIndex = input.index(after: currentIndex) } } // 用转换后的字节数组构建UTF-8字符串 return String(bytes: byteBuffer, encoding: .utf8) }
测试使用
let encodedSSID = "\\xc3\\xa6" if let decodedSSID = decodeEscapedUTF8(encodedSSID) { print(decodedSSID) // 输出:æ }
原理说明
- 转义序列识别:函数会遍历输入字符串,识别
\x开头的十六进制转义序列,将其转换成对应的UInt8字节。 - 字节构建与解码:收集所有转换后的字节(包括普通字符的UTF-8字节),最后用
String(bytes:encoding:)方法将字节数组解码成真正的UTF-8字符串。 - 容错处理:如果遇到不完整的转义序列(比如
\x后面不足两位十六进制数),函数会返回nil,你可以根据需求调整容错逻辑。
内容的提问来源于stack exchange,提问作者YoungChul
相关产品推荐
相关产品推荐

