Swift中解码JSON时如何去除HTML实体?
解决Swift中JSONDecoder处理HTML实体的问题
JSONDecoder本身没有内置处理HTML实体的功能,但可以通过自定义解码逻辑实现自动转换,以下是两种可靠的方案:
方案一:自定义HTML实体解码工具函数+结构体自定义解码
首先实现一个通用的HTML实体解码函数,利用NSAttributedString的HTML解析能力处理各种实体:
import Foundation func decodeHTMLEntities(_ string: String) -> String { guard let data = string.data(using: .utf8) else { return string } let options: [NSAttributedString.DocumentReadingOptionKey: Any] = [ .documentType: NSAttributedString.DocumentType.html, .characterEncoding: String.Encoding.utf8.rawValue ] do { let attributedString = try NSAttributedString(data: data, options: options, documentAttributes: nil) return attributedString.string } catch { print("HTML实体解码失败: \(error)") return string } }
然后修改Question结构体,自定义解码方法,在解码每个字符串字段时自动应用解码函数:
struct Question: Decodable { var type: String var difficulty: String var category: String var question: String var correctAnswer: String var incorrectAnswers: [String] enum CodingKeys: String, CodingKey { case type, difficulty, category, question case correctAnswer = "correct_answer" case incorrectAnswers = "incorrect_answers" } init(from decoder: Decoder) throws { let container = try decoder.container(keyedBy: CodingKeys.self) // 对单个字符串字段解码并处理实体 type = decodeHTMLEntities(try container.decode(String.self, forKey: .type)) difficulty = decodeHTMLEntities(try container.decode(String.self, forKey: .difficulty)) category = decodeHTMLEntities(try container.decode(String.self, forKey: .category)) question = decodeHTMLEntities(try container.decode(String.self, forKey: .question)) correctAnswer = decodeHTMLEntities(try container.decode(String.self, forKey: .correctAnswer)) // 对数组中的每个字符串处理实体 incorrectAnswers = try container.decode([String].self, forKey: .incorrectAnswers).map(decodeHTMLEntities) } }
修改后,原有的getQuestions方法无需改动,JSONDecoder会自动在解码Question时处理所有HTML实体,比如把'转换为',&转换为&。
方案二:自定义String的解码策略(全局生效)
如果希望所有字符串类型的字段都自动处理HTML实体,可以自定义String的扩展,替换默认解码逻辑:
import Foundation extension String: Decodable { public init(from decoder: Decoder) throws { let container = try decoder.singleValueContainer() let rawString = try container.decode(String.self) self = decodeHTMLEntities(rawString) } } // 复用之前的decodeHTMLEntities函数 func decodeHTMLEntities(_ string: String) -> String { guard let data = string.data(using: .utf8) else { return string } let options: [NSAttributedString.DocumentReadingOptionKey: Any] = [ .documentType: NSAttributedString.DocumentType.html, .characterEncoding: String.Encoding.utf8.rawValue ] do { let attributedString = try NSAttributedString(data: data, options: options, documentAttributes: nil) return attributedString.string } catch { print("HTML实体解码失败: \(error)") return string } }
注意:这种方法会全局替换所有String的解码行为,若项目中有不需要处理HTML实体的字符串字段,可能引发问题,因此更推荐方案一的针对性处理。
为什么不推荐解码后遍历处理?
虽然解码后遍历questions数组处理每个字段也能实现功能,但这种方式属于事后补救,代码冗余且容易遗漏字段。自定义解码逻辑是在解码过程中一次性处理,更符合Swift的类型安全和优雅编码原则。
内容的提问来源于stack exchange,提问作者John Citrowske
相关产品推荐
相关产品推荐

