如何修改正则表达式匹配跨行注释以获取完整Group3内容
解决方案
问题原因
你当前采用逐行匹配的逻辑,跨多行的续行注释仅能捕获第一行//后的内容,下一行带前置空白的//注释内容不会被同一次匹配纳入;同时原正则默认规则下.不会命中换行符,也没有处理续行注释的逻辑。
方案一:批量匹配完整文本(推荐)
不需要逐行读取,直接对完整待匹配文本做批量匹配,再格式化注释内容即可:
- 调整正则与匹配逻辑,开启单行模式让
.支持匹配换行,用非贪婪匹配避免跨条目匹配:
// yourFullText为完整的待匹配文本内容 Regex regexPattern = new Regex(@"<\w\w:Value> SYMBOL: (P.*?)=(.*?)//(.*?)(?=<\w\w:Value>|\z)", RegexOptions.Singleline); MatchCollection allEntries = regexPattern.Matches(yourFullText);
- 对捕获到的Group3内容做格式化,替换续行注释的标记为空格:
foreach (Match entry in allEntries) { string rawComment = entry.Groups[3].Value; // 替换所有换行+前置空白+//的片段为单个空格 string fullComment = Regex.Replace(rawComment, @"\r?\n\s*//\s*", " ").Trim(); // fullComment即为预期的完整注释内容 }
方案二:保留原有逐行读取逻辑
如果不想改动现有逐行读取的逻辑,可以新增临时变量拼接续行注释:
string unfinishedComment = string.Empty; Regex entryRegex = new Regex(@"<\w\w:Value> SYMBOL: (P.*)=(.*)//(.*)"); Regex continueCommentRegex = new Regex(@"^\s*//(.*)"); string line; // 假设你原有逻辑是用StreamReader逐行读取 while ((line = yourReader.ReadLine()) != null) { Match entryMatch = entryRegex.Match(line); if (entryMatch.Success) { // 先输出上一条的完整注释 if (!string.IsNullOrEmpty(unfinishedComment)) { Console.WriteLine(unfinishedComment.Trim()); } // 初始化当前条目的注释 unfinishedComment = entryMatch.Groups[3].Value.Trim() + " "; continue; } // 匹配到续行注释就拼接到当前注释 Match commentMatch = continueCommentRegex.Match(line); if (commentMatch.Success && !string.IsNullOrEmpty(unfinishedComment)) { unfinishedComment += commentMatch.Groups[1].Value.Trim() + " "; } } // 输出最后一条条目对应的注释 if (!string.IsNullOrEmpty(unfinishedComment)) { Console.WriteLine(unfinishedComment.Trim()); }
输出结果
两种方案最终得到的注释内容都符合你的预期:
Projektierung D-Weg Freimeldung nicht auswerten Länge des Durchrutschweges
内容的提问来源于stack exchange,提问作者user15519784
相关产品推荐
相关产品推荐

