Regex与strncpy未按预期工作,C语言代码问题求助
遇到的问题
- 处理包含多行多URL的文件时,正则匹配循环无限运行
- 处理单行单URL文件时,Valgrind提示
matches[count] = malloc(sizeof(char) * match_length);和matches[count][match_length] = 0;存在Invalid size 1 write错误,且sizeof(matches)/sizeof(matches[0])结果恒为1 - 误解
strncpy会自动在字符串末尾添加空字节 - 匹配结果包含换行符,需手动置0才能去除
问题原因与修复方案
1. 正则循环无限运行
原代码中do-while循环里,下一次regexec调用使用line + match.rm_eo + 1作为起始位置,跳过了当前匹配的结束字符,可能导致重复匹配同一区域,引发死循环。同时初始的if (regexec(...)) continue;逻辑冗余,易导致匹配流程混乱。
修复:
改用单while循环,每次将匹配起始偏移量更新为当前匹配的结束位置,确保每次匹配从上次结束处开始,避免重复或遗漏。
2. 内存越界错误
malloc(sizeof(char) * match_length)仅分配了match_length个字节,但matches[count][match_length] = 0;是往索引match_length的位置写入(数组索引从0开始,有效范围是0到match_length-1),属于越界访问。另外,matches是指针而非数组,sizeof(matches)/sizeof(matches[0])永远等于1(指针大小除以指针大小),不能用该表达式判断数组长度,必须依赖手动维护的count和size变量。
修复:
分配内存时预留空终止符位置:malloc(sizeof(char) * (match_length + 1)),确保写入空字节时不会越界。
3. strncpy的误解
strncpy的行为是:若源字符串长度小于指定长度,会用空字节填充至指定长度;若源字符串长度大于等于指定长度,仅复制指定长度的字符,不会自动添加空终止符。必须手动添加空字节以保证字符串合法。
修复:
复制完成后手动写入'\0',或使用strndup(系统支持时)自动处理空终止符。
4. 匹配结果含换行符
原正则表达式[[:alnum:][:punct:]]+会匹配到行尾的换行符(换行符不属于[:punct:],但实际匹配中可能因行尾字符边界问题被包含),导致结果带换行。
修复:
修正正则表达式,在末尾添加[^[:space:]],确保匹配到URL末尾时停止,排除空格、换行等空白字符。
修复后的完整代码
#include <stdio.h> #include <stdlib.h> #include <regex.h> #include <string.h> int extract_urls(FILE *file) { // 初始化动态匹配结果数组,初始容量为1 char **matches = malloc(sizeof(char *)); if (!matches) return 0; char *line = NULL; size_t line_capacity = 0; size_t count = 0, size = 1; ssize_t read_len; regex_t rgx; // 修正正则:排除空白字符,避免匹配换行符 int reg_err = regcomp(&rgx, "((gopher|https?)://[[:alnum:][:punct:]]+[^[:space:]])", REG_ICASE | REG_EXTENDED); if (reg_err != 0) { free(matches); return 0; } while ((read_len = getline(&line, &line_capacity, file)) != -1) { regmatch_t match; int offset = 0; int error; // 循环匹配当前行内所有URL while ((error = regexec(&rgx, line + offset, 1, &match, 0)) == 0) { int match_start = match.rm_so + offset; int match_end = match.rm_eo + offset; int match_length = match_end - match_start; // 分配含空终止符的内存空间 matches[count] = malloc(sizeof(char) * (match_length + 1)); if (!matches[count]) { // 内存分配失败,清理已分配资源 for (size_t i = 0; i < count; i++) free(matches[i]); free(matches); regfree(&rgx); free(line); return 0; } // 复制匹配内容并手动添加空终止符 strncpy(matches[count], line + match_start, match_length); matches[count][match_length] = '\0'; // 动态扩容数组 count++; if (count == size) { size *= 2; char **temp = realloc(matches, sizeof(char *) * size); if (!temp) { // 扩容失败,清理资源 for (size_t i = 0; i < count; i++) free(matches[i]); free(matches); regfree(&rgx); free(line); return 0; } matches = temp; } // 更新偏移量,避免重复匹配 offset = match_end; } } // 示例:打印所有匹配的URL并释放资源 for (size_t i = 0; i < count; i++) { printf("URL %zu: %s\n", i+1, matches[i]); free(matches[i]); } free(matches); regfree(&rgx); free(line); return 1; } // 测试主函数 int main() { FILE *file = fopen("urls.txt", "r"); if (!file) { perror("Failed to open file"); return 1; } extract_urls(file); fclose(file); return 0; }
内容的提问来源于stack exchange,提问作者tteixeira

