SAS EG中prxmatch匹配开头no且排除not误匹配的正则问题
问题根因
当前正则存在3个核心问题导致误过滤:
- 转义符冲突:SAS中双引号包裹的字符串会优先解析反斜杠转义,写的
\b会被转义为退格控制符,根本不会作为正则的单词边界元字符传递给PRX引擎,导致^no规则会直接匹配所有以no开头的字符串,包含开头为not的内容。 - 锚定规则缺失:正则或运算符
|的优先级最低,写的all good|na|n\/a|n\.a分支没有任何位置锚定,只要评论任意位置出现这些字符组合就会被命中,这是非开头位置的内容被误过滤的核心原因。 - 未兼容前导空白:真实采集的评论常带有前导空格、制表符等不可见字符,原正则的
^锚定只能匹配紧挨着字符串起始位置的关键词,会漏掉开头带空白的待过滤评论。
修正方案
- 用单引号包裹正则模式串,避免SAS转义反斜杠,保证正则元字符正常生效。
- 给所有过滤规则加上正确的位置锚定:
- 开头否定词(no、nil、none、nop)规则:匹配字符串起始位置+任意前导空白+目标关键词+单词边界,既保证仅匹配开头的关键词,又能通过
\b排除not这类no为单词前缀的场景。 - 无意义短评(all good、na、n/a、n.a)规则:匹配去掉前后空白后整个内容为目标短语的场景,避免误命中评论中间出现的对应字符。
- 修正后的完整代码如下:
data cclhd hnelhd islhd nbmlhd seslhd swslhd slhd wslhd mnclhd nnswlhd wnswlhd; set work.schools_dataset; where Comments ne " "; /* 修正后的PRX匹配规则 */ if prxmatch ('m/^\s*(no\b|nil|none|nop)|^\s*(all good|na|n\/a|n\.a)\s*$/i',Comments) = 0 ; keep ParticipantID FirstName Mobile VaxDate OperationID Venue Comments; if operationid=108 then output work.cclhd; else if operationid=109 then output work.hnelhd; else if operationid=110 then output work.islhd; else if operationid=111 then output work.nbmlhd; else if operationid=113 then output work.seslhd; else if operationid=114 then output work.swslhd; else if operationid=115 then output work.slhd; else if operationid=116 then output work.wslhd; else if operationid=118 then output work.mnclhd; else if operationid=120 then output work.nnswlhd; else if operationid=122 then output work.wnswlhd; run;
匹配效果验证
- 正常过滤(匹配命中)的内容:
No problems、No、Nil to report、NONE、all good、n/a、N.A - 正常保留(不命中)的内容:
His name is STEVE not SPEVE、I was mostly fine but I was not expecting to get a headache、Not any side effects、I have no discomfort today
内容的提问来源于stack exchange,提问作者Renee
相关产品推荐
相关产品推荐

