在C语言中使用POSIX正则函数获取正则捕获组的问题
问题背景
想要用C语言的<regex.h>库模仿PostgreSQL的regexp_match函数,以下是参考的两条PostgreSQL查询:
SELECT regexp_match('hello world test', 'world.{3}'); SELECT regexp_match('foobarbequebaz', '(bar)(beque)');
测试时发现,第二条查询对应的C代码本应输出两个捕获组(group1仅返回bar),但出现异常;后续验证pattern2匹配sz2时行为正常,但对无捕获组模式下的matches数组存在疑问。
测试代码
#include<regex.h> #include<stdio.h> #include<string.h> #include<stdlib.h> #define MAX_MATCHES 1024 int main(void) { regex_t regex; int reti; char msgbuf[100]; char buff0[20]; char buff[20]; char buff1[20]; char *sz1 = "hello world test"; //char *sz2= "foobarbequebaz"; char *pattern1 = "world.{3}"; //char *pattern2 = "(bar)(beque)"; regmatch_t matches[MAX_MATCHES]; /* Compile regular expression */ reti = regcomp(®ex,pattern1,REG_EXTENDED); if(reti){ fprintf(stderr,"could not compile\n"); exit(EXIT_FAILURE); } reti = regexec(®ex,sz1,MAX_MATCHES,matches,0); if(!reti){ printf("szso=%d\n",matches[1].rm_so); printf("szeo=%d\n",matches[1].rm_eo); memcpy(buff0,sz1+matches[0].rm_so,matches[0].rm_eo-matches[0].rm_so); memcpy(buff,sz1+matches[1].rm_so,matches[1].rm_eo-matches[1].rm_so); memcpy(buff1,sz1+matches[2].rm_so,matches[2].rm_eo-matches[2].rm_so); printf("group0: %s\n",buff0); printf("group1: %s\n",buff); printf("group2: %s\n",buff1); } else if(reti == REG_NOMATCH){ puts("No match"); } else{ regerror(reti,®ex,msgbuf,sizeof(msgbuf)); fprintf(stderr,"Regex match failed: %s\n",msgbuf); exit(EXIT_FAILURE); } regfree(®ex); exit(EXIT_SUCCESS); }
异常输出(测试pattern1时)
szso=3 szeo=11 group1: barbeque
核心疑问
- 当模式无捕获组时,
matches[0]是否应与matches[1]内容相同? - 此类场景下,
group0和group1是否应该一致?
问题分析与解决方案
1. 捕获组的规则
POSIX正则中,matches数组的第0项固定对应整个正则匹配的内容;从第1项开始,才对应正则中定义的捕获组。如果正则没有定义任何捕获组(比如pattern1: world.{3}),那么matches[1]及之后的项是未定义的,访问它们会读取到内存中的随机值,这就是异常输出的原因。
2. 获取有效捕获组数量
编译正则后,可以通过regex.re_nsub变量获取正则中定义的捕获组总数。比如pattern2: (bar)(beque)对应的regex.re_nsub值为2,此时matches[1]和matches[2]才是有效的捕获组内容。
3. 代码修正关键点
- 访问
matches数组时,不要超过regex.re_nsub的范围(加上第0项) - 使用
memcpy截取字符串后,必须手动添加字符串结束符\0,否则打印时会输出垃圾内容(原字符串的截取部分默认无结束符)
修正后的示例代码
#include<regex.h> #include<stdio.h> #include<string.h> #include<stdlib.h> #define MAX_MATCHES 1024 int main(void) { regex_t regex; int reti; char msgbuf[100]; char buff0[20] = {0}; // 初始化内存避免垃圾值 char buff[20] = {0}; char buff1[20] = {0}; char *sz2= "foobarbequebaz"; char *pattern2 = "(bar)(beque)"; regmatch_t matches[MAX_MATCHES]; reti = regcomp(®ex, pattern2, REG_EXTENDED); if(reti){ fprintf(stderr,"could not compile\n"); exit(EXIT_FAILURE); } reti = regexec(®ex, sz2, MAX_MATCHES, matches, 0); if(!reti){ // 处理整个匹配结果(group0) int len0 = matches[0].rm_eo - matches[0].rm_so; if(len0 < sizeof(buff0)){ memcpy(buff0, sz2 + matches[0].rm_so, len0); buff0[len0] = '\0'; // 添加字符串结束符 } printf("group0: %s\n", buff0); // 遍历所有有效捕获组 for(int i=1; i<=regex.re_nsub; i++){ int len = matches[i].rm_eo - matches[i].rm_so; char *buf = NULL; switch(i){ case 1: buf = buff; break; case 2: buf = buff1; break; default: continue; } if(len < sizeof(buf)){ memcpy(buf, sz2 + matches[i].rm_so, len); buf[len] = '\0'; } printf("group%d: %s\n", i, buf); } } else if(reti == REG_NOMATCH){ puts("No match"); } else{ regerror(reti,®ex,msgbuf,sizeof(msgbuf)); fprintf(stderr,"Regex match failed: %s\n",msgbuf); exit(EXIT_FAILURE); } regfree(®ex); exit(EXIT_SUCCESS); }
疑问解答
- 当模式无捕获组时,
matches[1]及之后的项是未定义的,和matches[0]没有关联,不能期望内容一致。只有正则中定义了捕获组,对应的matches项才有效。 - 只有当捕获组的范围完全覆盖整个正则匹配结果时(比如正则为
^(bar)$),matches[0]和对应的捕获组内容才会相同,其他场景下两者内容不一致。
内容的提问来源于stack exchange,提问作者jian
相关产品推荐
相关产品推荐

