使用|运算符合并NSRegularExpression后,如何识别匹配的具体模式?
解决NSRegularExpression合并正则后无法区分加粗/斜体的问题
嘿,这个问题我之前也碰到过!合并正则后分不清匹配的是加粗还是斜体,本质是因为合并后的正则包含了多个捕获组,我们只需要判断哪个捕获组对应有效内容,或者用更清晰的命名捕获组来区分,就能一次遍历搞定两种格式,不用多次调用检测函数啦。
方案一:通过捕获组索引判断类型
合并后的正则\*{2}([\w ]+)\*{2}|\_{1}([\w ]+)\_{1}有两个捕获组:第一个对应加粗的内容,第二个对应斜体的内容。匹配时,只有当前模式对应的捕获组会有有效范围,另一个则是NSRange(location: NSNotFound, length: 0)。我们可以利用这一点来判断格式类型:
let textView = UITextView() textView.text = "这是**加粗内容**,这是_斜体内容_" let pattern = "\\*{2}([\\w ]+)\\*{2}|\\_{1}([\\w ]+)\\_{1}" do { let regex = try NSRegularExpression(pattern: pattern, options: []) let fullRange = NSRange(location: 0, length: textView.text.utf16.count) let matches = regex.matches(in: textView.text, options: [], range: fullRange) for match in matches { // 检查第一个捕获组(加粗内容) let boldContentRange = match.range(at: 1) if boldContentRange.location != NSNotFound { let attributes: [NSAttributedString.Key: Any] = [ .font: UIFont.boldSystemFont(ofSize: textView.font?.pointSize ?? 17) ] textView.textStorage.addAttributes(attributes, range: boldContentRange) } // 检查第二个捕获组(斜体内容) let italicContentRange = match.range(at: 2) if italicContentRange.location != NSNotFound { let attributes: [NSAttributedString.Key: Any] = [ .font: UIFont.italicSystemFont(ofSize: textView.font?.pointSize ?? 17) ] textView.textStorage.addAttributes(attributes, range: italicContentRange) } } } catch { print("正则初始化失败:\(error.localizedDescription)") }
方案二:使用命名捕获组(可读性更高)
如果觉得靠索引判断容易搞混,推荐给每个捕获组起个名字,这样代码可读性更强,后期维护也更方便。修改正则为带命名捕获组的格式:
let pattern = "\\*{2}(?<boldContent>[\\w ]+)\\*{2}|\\_{1}(?<italicContent>[\\w ]+)\\_{1}" do { let regex = try NSRegularExpression(pattern: pattern, options: []) let fullRange = NSRange(location: 0, length: textView.text.utf16.count) let matches = regex.matches(in: textView.text, options: [], range: fullRange) for match in matches { // 通过名字获取加粗内容的范围 if let boldRange = regex.range(withName: "boldContent", in: match), boldRange.location != NSNotFound { let attributes: [NSAttributedString.Key: Any] = [ .font: UIFont.boldSystemFont(ofSize: textView.font?.pointSize ?? 17) ] textView.textStorage.addAttributes(attributes, range: boldRange) } // 通过名字获取斜体内容的范围 if let italicRange = regex.range(withName: "italicContent", in: match), italicRange.location != NSNotFound { let attributes: [NSAttributedString.Key: Any] = [ .font: UIFont.italicSystemFont(ofSize: textView.font?.pointSize ?? 17) ] textView.textStorage.addAttributes(attributes, range: italicRange) } } } catch { print("正则初始化失败:\(error.localizedDescription)") }
额外优化:完善正则匹配范围
你原来的正则用[\w ]只能匹配字母、数字、下划线和空格,如果文本里有标点、特殊字符(比如中文)就会匹配失败。建议把匹配内容的部分改成排除分隔符的写法:
- 加粗:用
[^*]+代替[\w ]+,表示匹配除了*之外的所有字符(避免中间出现**导致匹配错误) - 斜体:用
[^_]+代替[\w ]+,同理避免中间出现_
优化后的正则:
\*{2}([^*]+)\*{2}|\_{1}([^_]+)\_{1}
(命名捕获组版本对应改成\*{2}(?<boldContent>[^*]+)\*{2}|\\_{1}(?<italicContent>[^_]+)\_{1})
这样就能兼容更多场景的文本内容啦!
内容的提问来源于stack exchange,提问作者Ivan Cantarino
相关产品推荐
相关产品推荐

