在ClojureScript中如何无需环视按正则拆分字符串并保留匹配项
在ClojureScript中按正则拆分并保留匹配结果
因为正则环视在部分浏览器中支持不佳,我们可以换一种手动分割的思路:先定位所有匹配项的位置,再逐个分割字符串并把匹配项插入结果列表。
实现函数
(defn split-with-matches [s re] (let [matches (->> (re-seq re s) (map #(vector % (.indexOf s %))) (map (fn [[match start]] [match start (+ start (count match))]))) initial-start 0] (loop [matches matches prev-end initial-start result []] (if (empty? matches) (if (< prev-end (count s)) (conj result (subs s prev-end)) result) (let [[match start end] (first matches)] (recur (rest matches) end (-> result (cond-> (> start prev-end) (conj (subs s prev-end start))) (conj match))))))))
测试示例
;; 基础场景 (split-with-matches "foo {{bar}} baz" #"\{\{\s*[A-Za-z0-9_\.]+?\s*\}\}") ;; => ("foo " "{{bar}}" " baz") ;; 匹配项在开头 (split-with-matches "{{foo}} bar" #"\{\{\s*[A-Za-z0-9_\.]+?\s*\}\}") ;; => ("{{foo}}" " bar") ;; 匹配项在结尾 (split-with-matches "bar {{foo}}" #"\{\{\s*[A-Za-z0-9_\.]+?\s*\}\}") ;; => ("bar " "{{foo}}") ;; 无匹配项 (split-with-matches "no matches here" #"\{\{\s*[A-Za-z0-9_\.]+?\s*\}\}") ;; => ("no matches here")
逻辑说明
- 先用
re-seq获取所有匹配的子串,再通过.indexOf定位每个子串的起始索引,计算出结束索引 - 用
loop迭代所有匹配项,维护上一个匹配的结束位置:- 每次处理匹配项时,先把上一个结束位置到当前匹配起始位置的非匹配子串加入结果(如果存在)
- 再把当前匹配的子串加入结果
- 迭代结束后,处理字符串剩余的部分(如果有)
这个方案只依赖ClojureScript的基础API,不需要环视语法,能兼容所有目标浏览器。
内容的提问来源于stack exchange,提问作者noahmoss
相关产品推荐
相关产品推荐

