You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Swift中文本字符串轻量化清理方案技术问询

轻量化Swift文本清理方案

嘿,针对你遇到的这种来源不可控的杂乱文本清理需求,我刚好有几个轻量的Swift方案,完全不用依赖第三方重型库,靠原生API就能搞定常见的格式问题!

核心清理步骤拆解

1. 移除HTML标签

像你例子里的<p>这类标签,用正则就能快速剥离:

func stripHTMLTags(from text: String) -> String {
    return text.replacingOccurrences(of: "<[^>]+>", with: "", options: .regularExpression)
}

2. 替换非断空格( )

网页文本里常见的非断空格,不管是实体形式的&nbsp;还是Unicode编码的\u{00A0},直接替换成普通空格就行:

func replaceNonBreakingSpaces(in text: String) -> String {
    return text.replacingOccurrences(of: "\u{00A0}", with: " ")
               .replacingOccurrences(of: "&nbsp;", with: " ")
}

3. 清理自定义强调标记(\emphasize\)

针对你例子里的反斜杠包裹标记,用正则匹配并移除多余的反斜杠:

func stripCustomEmphasisMarkers(from text: String) -> String {
    // 匹配\内容\的格式,捕获中间文本并替换
    return text.replacingOccurrences(of: #"\\(.*?)\\"#, with: "$1", options: .regularExpression)
}

整合所有清理逻辑

把上面的步骤打包成一个统一函数,一次处理所有问题,还能额外清理多余空格和首尾空白:

func cleanRawText(_ rawText: String) -> String {
    var cleanedText = rawText
    cleanedText = stripHTMLTags(from: cleanedText)
    cleanedText = replaceNonBreakingSpaces(in: cleanedText)
    cleanedText = stripCustomEmphasisMarkers(from: cleanedText)
    // 可选:合并连续空格
    cleanedText = cleanedText.replacingOccurrences(of: "\\s+", with: " ", options: .regularExpression)
    // 去除首尾空白
    return cleanedText.trimmingCharacters(in: .whitespacesAndNewlines)
}

测试你的示例文本

用你给出的两个测试用例验证效果:

// 第一个测试用例
let rawText1 = "<p>Occasionally we&nbsp;deal&nbsp;with this.</p>"
print(cleanRawText(rawText1)) // 输出: "Occasionally we deal with this."

// 第二个测试用例
let rawText2 = "<p>Sometimes they \\emphasize\\ like this, I could live with it</p>"
print(cleanRawText(rawText2)) // 输出: "Sometimes they emphasize like this, I could live with it"

这个方案完全基于Swift原生能力,轻量化易扩展,后续遇到新的格式问题,直接往函数里加新的处理步骤就行~

内容的提问来源于stack exchange,提问作者Dale

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:26:02