Swift中文本字符串轻量化清理方案技术问询
轻量化Swift文本清理方案
嘿,针对你遇到的这种来源不可控的杂乱文本清理需求,我刚好有几个轻量的Swift方案,完全不用依赖第三方重型库,靠原生API就能搞定常见的格式问题!
核心清理步骤拆解
1. 移除HTML标签
像你例子里的<p>这类标签,用正则就能快速剥离:
func stripHTMLTags(from text: String) -> String { return text.replacingOccurrences(of: "<[^>]+>", with: "", options: .regularExpression) }
2. 替换非断空格( )
网页文本里常见的非断空格,不管是实体形式的 还是Unicode编码的\u{00A0},直接替换成普通空格就行:
func replaceNonBreakingSpaces(in text: String) -> String { return text.replacingOccurrences(of: "\u{00A0}", with: " ") .replacingOccurrences(of: " ", with: " ") }
3. 清理自定义强调标记(\emphasize\)
针对你例子里的反斜杠包裹标记,用正则匹配并移除多余的反斜杠:
func stripCustomEmphasisMarkers(from text: String) -> String { // 匹配\内容\的格式,捕获中间文本并替换 return text.replacingOccurrences(of: #"\\(.*?)\\"#, with: "$1", options: .regularExpression) }
整合所有清理逻辑
把上面的步骤打包成一个统一函数,一次处理所有问题,还能额外清理多余空格和首尾空白:
func cleanRawText(_ rawText: String) -> String { var cleanedText = rawText cleanedText = stripHTMLTags(from: cleanedText) cleanedText = replaceNonBreakingSpaces(in: cleanedText) cleanedText = stripCustomEmphasisMarkers(from: cleanedText) // 可选:合并连续空格 cleanedText = cleanedText.replacingOccurrences(of: "\\s+", with: " ", options: .regularExpression) // 去除首尾空白 return cleanedText.trimmingCharacters(in: .whitespacesAndNewlines) }
测试你的示例文本
用你给出的两个测试用例验证效果:
// 第一个测试用例 let rawText1 = "<p>Occasionally we deal with this.</p>" print(cleanRawText(rawText1)) // 输出: "Occasionally we deal with this." // 第二个测试用例 let rawText2 = "<p>Sometimes they \\emphasize\\ like this, I could live with it</p>" print(cleanRawText(rawText2)) // 输出: "Sometimes they emphasize like this, I could live with it"
这个方案完全基于Swift原生能力,轻量化易扩展,后续遇到新的格式问题,直接往函数里加新的处理步骤就行~
内容的提问来源于stack exchange,提问作者Dale
相关产品推荐
相关产品推荐

